- Corpus: 275 passages
- Queries: 110 labeled queries
- Ranking: exact dense dot-product retrieval over the same corpus
- Prefixes:
query:for queries,passage:for corpus passages - Models are loaded in separate processes on the same GPU.
Quality + performance
| Variant | Dim | Hit@1 | Hit@5 | Hit@10 | Recall@10 | MRR@10 | nDCG@10 | MAP@10 | Median rank | Hard-neg win | Docs/s | Queries/s | Peak Δ VRAM GiB | 1M vectors f32 GiB |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1B-AWQ / native_2048d | 2048 | 91.82% | 100.00% | 100.00% | 100.00% | 0.9545 | 0.9662 | 0.9545 | 1.0 | 94.55% | 1003.8 | 1254.9 | 2.32 | 7.63 |
| 8B-AWQ / native_4096d | 4096 | 92.73% | 100.00% | 100.00% | 100.00% | 0.9621 | 0.9720 | 0.9621 | 1.0 | 94.55% | 181.0 | 239.9 | 7.35 | 15.26 |
| 8B-AWQ / sliced_2048d | 2048 | 94.55% | 100.00% | 100.00% | 100.00% | 0.9712 | 0.9787 | 0.9712 | 1.0 | 96.36% | 181.0 | 239.9 | 7.35 | 7.63 |
Paired bootstrap
All deltas below are A - B on the same queries. A 95% CI crossing zero means the benchmark did not establish a clear direction for that metric.
8B/sliced_2048d vs 1B/native_2048d
First-relevant-rank wins/ties/losses for A: 4 / 104 / 2
| Metric | Δ A-B | 95% bootstrap CI | bootstrap P(Δ>0) |
|---|---|---|---|
| hit@1 | +0.02727 | [-0.00909, +0.07273] | 87.5% |
| hit@5 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| hit@10 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| recall@10 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| mrr@10 | +0.01667 | [-0.00606, +0.04318] | 91.7% |
| ndcg@10 | +0.01252 | [-0.00455, +0.03259] | 91.9% |
| ap@10 | +0.01667 | [-0.00606, +0.04318] | 91.7% |
8B/native_4096d vs 1B/native_2048d
First-relevant-rank wins/ties/losses for A: 4 / 102 / 4
| Metric | Δ A-B | 95% bootstrap CI | bootstrap P(Δ>0) |
|---|---|---|---|
| hit@1 | +0.00909 | [-0.03636, +0.05455] | 57.5% |
| hit@5 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| hit@10 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| recall@10 | +0.00000 | [+0.00000, +0.00000] | 0.0% |
| mrr@10 | +0.00758 | [-0.01894, +0.03712] | 68.9% |
| ndcg@10 | +0.00581 | [-0.01398, +0.02776] | 70.5% |
| ap@10 | +0.00758 | [-0.01894, +0.03712] | 68.9% |
How to interpret this
- Compare 8B sliced 2048D vs 1B native 2048D first. That is the storage-matched quality comparison.
- Then compare 8B native 4096D vs 1B native 2048D. That answers whether paying for the 8B’s full vector size improves your actual retrieval.
- Do not choose the 8B from a tiny MRR difference if its confidence interval crosses zero. Inspect tag-level results and the queries where rankings materially changed.
- If your MemPalace is bilingual, include Spanish→English and English→Spanish queries, temporal references, paraphrases, vague references, and near-duplicate memories.
Failure inspection
Detailed top-5 retrievals for the strongest disagreements are written to failure_examples.json.