Back to the experimentSupporting notebook

Nemotron 1B vs 8B AWQ for MemPalace retrieval (2026-09-18)

Source: nemotron/embed-1b-vs-8b-awq-2026-09-18/README.md · revision 6550ead3945b

Plain-language summary: FINDINGS.md — winners, why they make sense, caveats, deployed choice.

Question: for MemPalace bilingual (EN/ES) memory retrieval, is the 8B AWQ encoder better than the 1B AWQ — as a better encoder at matched 2048D storage, and at its full 4096D size — and what does it cost in throughput, VRAM, and index size?

1. Dataset: synthetic MemPalace-like memory bench

Built by scripts/build_dataset.py: 55 hand-written topic clusters (benchmarks, bugs, infra, business, policy, preferences), each with 1 gold passage + 2 hard negatives (same topic, different key fact) + 2 fillers, and 2 labeled queries (typically 1 EN + 1 ES).

2. Measured results (exact dense dot-product retrieval, same corpus)

Full tables: data/REPORT.md · machine-readable: data/summary.json · per-variant detail: data/1b_result.json, data/8b_result.json

Variant Dim Hit@1 MRR@10 nDCG@10 Hard-neg win Docs/s Queries/s Peak ΔVRAM 1M vec f32
1B native 2048 91.82% 0.9545 0.9662 94.55% 1003.8 1254.9 2.32 GiB 7.63 GiB
8B native 4096 92.73% 0.9621 0.9720 94.55% 181.0 239.9 7.35 GiB 15.26 GiB
8B sliced+renorm 2048 94.55% 0.9712 0.9787 96.36% 181.0 239.9 7.35 GiB 7.63 GiB

Model load: 1B 88.1 s, 8B 122.3 s (includes HF download on first run). All variants reach 100% Hit@5/10/20. Worker logs: data/1B-AWQ.worker.log, data/8B-AWQ.worker.log.

Paired bootstrap (5000 samples, A−B on the same 110 queries)

detailed CIs · strongest disagreements

Comparison ΔHit@1 [95% CI] ΔMRR@10 [95% CI] Rank wins/ties/losses (A) P(Δ>0) MRR
8B/sliced-2048D vs 1B/native +0.027 [−0.009, +0.073] +0.017 [−0.006, +0.043] 4 / 104 / 2 91.7%
8B/native-4096D vs 1B/native +0.009 [−0.036, +0.055] +0.008 [−0.019, +0.037] 4 / 102 / 4 68.9%

Both 95% CIs cross zero: no statistically established quality win for the 8B on this dataset. The storage-matched slice leans 8B (92% bootstrap probability) but is unproven.

Tag slices (n=110; small-tag cells are indicative only)

Tag n 1B MRR 8B-native MRR 8B-sliced MRR
cross-lingual 55 0.9500 0.9515 0.9606
en 55 0.9500 0.9606 0.9697
es 55 0.9591 0.9636 0.9727
paraphrase 59 0.9703 0.9746 0.9831
technical 73 0.9795 0.9795 0.9932
temporal 10 0.9500 0.8500 0.9500
vague-reference 3 0.8333 1.0000 1.0000

3. Environment notes (replication hazards found)

  1. pip install "vllm==0.25.0" numpy alone is not sufficient: the worker import chain needs scipy (via transformers) and threadpoolctl + msgpack (via sklearn/librosa). All pinned in environment.yml.
  2. Latest transformers (5.17.0 at run time) breaks from vllm import LLM with a masked GenerationMixin import error; 5.12.1 works (matches the known-good qwen38-orca env).
  3. vLLM uses lazy imports, so import vllm; vllm.__version__ succeeding proves nothing — verify with from vllm import LLM.

4. Interpretation (opinion, not measurement)

5. Limits

6. Reproduce

conda env create -f scripts/environment.yml
bash scripts/run_parallel.sh 0 1   # 1B on GPU 0, 8B on GPU 1

Needs ~7 GB free for the two model downloads plus the conda env, and 2 free GPUs (or pass the same GPU twice / --sequential for one card).