Back to the experimentSupporting notebook

Nemotron MemPalace benchmark

Source: nemotron/embed-1b-vs-8b-awq-2026-09-18/data/REPORT.md · revision 6550ead3945b

Quality + performance

Variant Dim Hit@1 Hit@5 Hit@10 Recall@10 MRR@10 nDCG@10 MAP@10 Median rank Hard-neg win Docs/s Queries/s Peak Δ VRAM GiB 1M vectors f32 GiB
1B-AWQ / native_2048d 2048 91.82% 100.00% 100.00% 100.00% 0.9545 0.9662 0.9545 1.0 94.55% 1003.8 1254.9 2.32 7.63
8B-AWQ / native_4096d 4096 92.73% 100.00% 100.00% 100.00% 0.9621 0.9720 0.9621 1.0 94.55% 181.0 239.9 7.35 15.26
8B-AWQ / sliced_2048d 2048 94.55% 100.00% 100.00% 100.00% 0.9712 0.9787 0.9712 1.0 96.36% 181.0 239.9 7.35 7.63

Paired bootstrap

All deltas below are A - B on the same queries. A 95% CI crossing zero means the benchmark did not establish a clear direction for that metric.

8B/sliced_2048d vs 1B/native_2048d

First-relevant-rank wins/ties/losses for A: 4 / 104 / 2

Metric Δ A-B 95% bootstrap CI bootstrap P(Δ>0)
hit@1 +0.02727 [-0.00909, +0.07273] 87.5%
hit@5 +0.00000 [+0.00000, +0.00000] 0.0%
hit@10 +0.00000 [+0.00000, +0.00000] 0.0%
recall@10 +0.00000 [+0.00000, +0.00000] 0.0%
mrr@10 +0.01667 [-0.00606, +0.04318] 91.7%
ndcg@10 +0.01252 [-0.00455, +0.03259] 91.9%
ap@10 +0.01667 [-0.00606, +0.04318] 91.7%

8B/native_4096d vs 1B/native_2048d

First-relevant-rank wins/ties/losses for A: 4 / 102 / 4

Metric Δ A-B 95% bootstrap CI bootstrap P(Δ>0)
hit@1 +0.00909 [-0.03636, +0.05455] 57.5%
hit@5 +0.00000 [+0.00000, +0.00000] 0.0%
hit@10 +0.00000 [+0.00000, +0.00000] 0.0%
recall@10 +0.00000 [+0.00000, +0.00000] 0.0%
mrr@10 +0.00758 [-0.01894, +0.03712] 68.9%
ndcg@10 +0.00581 [-0.01398, +0.02776] 70.5%
ap@10 +0.00758 [-0.01894, +0.03712] 68.9%

How to interpret this

  1. Compare 8B sliced 2048D vs 1B native 2048D first. That is the storage-matched quality comparison.
  2. Then compare 8B native 4096D vs 1B native 2048D. That answers whether paying for the 8B’s full vector size improves your actual retrieval.
  3. Do not choose the 8B from a tiny MRR difference if its confidence interval crosses zero. Inspect tag-level results and the queries where rankings materially changed.
  4. If your MemPalace is bilingual, include Spanish→English and English→Spanish queries, temporal references, paraphrases, vague references, and near-duplicate memories.

Failure inspection

Detailed top-5 retrievals for the strongest disagreements are written to failure_examples.json.