Plain-language summary: FINDINGS.md — winners, why they make sense, caveats, deployed choice.
- Date: 2026-09-18 (single session: env, dataset, parallel run, analysis)
- Host: Ubuntu 24.04.5 LTS, 3x RTX 3090 24 GB (driver 595.58.03), 20 CPUs, 128 GB RAM
- Models:
groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16(rev452afdda, 1.04 GB safetensors) → GPU 0groxaxo/Nemotron-3-Embed-8B-AWQ-W4A16-G32-ASYM(rev58055100, 5.36 GB safetensors) → GPU 1
- Stack: conda env
nemotron-bench(python 3.12.14, vLLM 0.25.0, transformers 5.12.1, torch 2.11+cu130, numpy 2.3.5, scipy 1.18.1) — see scripts/environment.yml - Run mode: both workers in parallel, 1 GPU each, separate processes pinned via
CUDA_VISIBLE_DEVICES; rerun with scripts/run_parallel.sh
Question: for MemPalace bilingual (EN/ES) memory retrieval, is the 8B AWQ encoder better than the 1B AWQ — as a better encoder at matched 2048D storage, and at its full 4096D size — and what does it cost in throughput, VRAM, and index size?
1. Dataset: synthetic MemPalace-like memory bench
Built by scripts/build_dataset.py: 55 hand-written topic clusters (benchmarks, bugs, infra, business, policy, preferences), each with 1 gold passage + 2 hard negatives (same topic, different key fact) + 2 fillers, and 2 labeled queries (typically 1 EN + 1 ES).
- Corpus: 275 passages (data/mempalace_corpus.jsonl)
- Queries: 110 labeled, every query carries 2 hard negatives (data/mempalace_queries.jsonl)
- Validation: data/dataset_validation.json — no warnings, zero lexical-leakage flags (Jaccard ≥ 0.65)
- Prefixes:
query:/passage:, applied identically to both models
2. Measured results (exact dense dot-product retrieval, same corpus)
Full tables: data/REPORT.md · machine-readable: data/summary.json · per-variant detail: data/1b_result.json, data/8b_result.json
| Variant | Dim | Hit@1 | MRR@10 | nDCG@10 | Hard-neg win | Docs/s | Queries/s | Peak ΔVRAM | 1M vec f32 |
|---|---|---|---|---|---|---|---|---|---|
| 1B native | 2048 | 91.82% | 0.9545 | 0.9662 | 94.55% | 1003.8 | 1254.9 | 2.32 GiB | 7.63 GiB |
| 8B native | 4096 | 92.73% | 0.9621 | 0.9720 | 94.55% | 181.0 | 239.9 | 7.35 GiB | 15.26 GiB |
| 8B sliced+renorm | 2048 | 94.55% | 0.9712 | 0.9787 | 96.36% | 181.0 | 239.9 | 7.35 GiB | 7.63 GiB |
Model load: 1B 88.1 s, 8B 122.3 s (includes HF download on first run). All variants reach 100% Hit@5/10/20. Worker logs: data/1B-AWQ.worker.log, data/8B-AWQ.worker.log.
Paired bootstrap (5000 samples, A−B on the same 110 queries)
detailed CIs · strongest disagreements
| Comparison | ΔHit@1 [95% CI] | ΔMRR@10 [95% CI] | Rank wins/ties/losses (A) | P(Δ>0) MRR |
|---|---|---|---|---|
| 8B/sliced-2048D vs 1B/native | +0.027 [−0.009, +0.073] | +0.017 [−0.006, +0.043] | 4 / 104 / 2 | 91.7% |
| 8B/native-4096D vs 1B/native | +0.009 [−0.036, +0.055] | +0.008 [−0.019, +0.037] | 4 / 102 / 4 | 68.9% |
Both 95% CIs cross zero: no statistically established quality win for the 8B on this dataset. The storage-matched slice leans 8B (92% bootstrap probability) but is unproven.
Tag slices (n=110; small-tag cells are indicative only)
| Tag | n | 1B MRR | 8B-native MRR | 8B-sliced MRR |
|---|---|---|---|---|
| cross-lingual | 55 | 0.9500 | 0.9515 | 0.9606 |
| en | 55 | 0.9500 | 0.9606 | 0.9697 |
| es | 55 | 0.9591 | 0.9636 | 0.9727 |
| paraphrase | 59 | 0.9703 | 0.9746 | 0.9831 |
| technical | 73 | 0.9795 | 0.9795 | 0.9932 |
| temporal | 10 | 0.9500 | 0.8500 | 0.9500 |
| vague-reference | 3 | 0.8333 | 1.0000 | 1.0000 |
3. Environment notes (replication hazards found)
pip install "vllm==0.25.0" numpyalone is not sufficient: the worker import chain needsscipy(via transformers) andthreadpoolctl+msgpack(via sklearn/librosa). All pinned inenvironment.yml.- Latest
transformers(5.17.0 at run time) breaksfrom vllm import LLMwith a maskedGenerationMixinimport error; 5.12.1 works (matches the known-goodqwen38-orcaenv). - vLLM uses lazy imports, so
import vllm; vllm.__version__succeeding proves nothing — verify withfrom vllm import LLM.
4. Interpretation (opinion, not measurement)
- The 1B is the pragmatic MemPalace encoder: quality within noise of the 8B while embedding 5.5x faster (1004 vs 181 docs/s), using 3x less VRAM, and halving index storage at native dims.
- The 8B’s full 4096D buys nothing measurable here — its own 2048D slice scored higher, suggesting the extra dims add noise rather than signal for short memory passages (hypothesis only; needs a harder dataset to confirm).
- To give the 8B a real chance to separate: re-run on harder data (real MemPalace queries, near-duplicate memories) where Hit@5 < 100%. The harness supports this unchanged — swap the two JSONL files.
5. Limits
- Synthetic dataset, not real MemPalace traffic; absolute scores do not transfer.
- n=110 queries: powered for large effects only; the observed small deltas are correctly reported as inconclusive.
- Short passages (1–2 sentences); long-document retrieval untested.
- Single run per model (no seed/infra variance); determinism comes from exact retrieval + fixed tie-breaks, not repeated trials.
- Throughput measured on RTX 3090 with batch 32,
gpu_memory_utilization=0.90,max_model_len=4096, bf16 pooling runner — comparison is apples-to-apples, but absolute numbers are config-specific.
6. Reproduce
conda env create -f scripts/environment.yml
bash scripts/run_parallel.sh 0 1 # 1B on GPU 0, 8B on GPU 1
Needs ~7 GB free for the two model downloads plus the conda env, and 2 free GPUs
(or pass the same GPU twice / --sequential for one card).