Full evidence: README.md · raw data: data/ · repro: scripts/
Winner: the 1B (pragmatic choice)
On 110 labeled bilingual memory queries over 275 passages, the 8B shows no statistically proven quality lead — both 95% bootstrap CIs cross zero — while the 1B embeds 5.5x faster (1004 vs 181 docs/s), uses ~3x less VRAM (2.3 vs 7.4 GiB), and halves index size (7.6 vs 15.3 GiB per 1M vectors at native dims).
Why it makes sense
- Short memory passages (1–2 sentences) don’t need 8B capacity; both models saturate at 100% Hit@5, so the 8B has nowhere to show an advantage.
- The 8B’s full 4096D vector scored below its own truncated 2048D slice (MRR 0.9621 vs 0.9712) — the extra dimensions look like noise for this workload.
- The storage-matched comparison (8B@2048D vs 1B@2048D) leans 8B at 92% bootstrap probability, but “leans” is not “proven”.
Caveats
- Synthetic dataset, not real MemPalace traffic; too easy (100% Hit@5+) to separate the models. A harder re-run on real queries could change the picture.
- n=110 only detects large effects; small-tag slices (temporal n=10, vague n=3) are anecdotes, not conclusions.
- One run per model; long-document retrieval untested.
Deployed config
1B-AWQ (groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16) as the MemPalace encoder:
native 2048D, query: /passage: prefixes, L2-normalized cosine retrieval.
Revisit only if real-traffic eval shows Hit@1 gaps the 8B demonstrably closes.