Back to the experimentSupporting notebook

Findings: Nemotron 1B vs 8B AWQ for MemPalace retrieval

Source: nemotron/embed-1b-vs-8b-awq-2026-09-18/FINDINGS.md · revision 6550ead3945b

Full evidence: README.md · raw data: data/ · repro: scripts/

Winner: the 1B (pragmatic choice)

On 110 labeled bilingual memory queries over 275 passages, the 8B shows no statistically proven quality lead — both 95% bootstrap CIs cross zero — while the 1B embeds 5.5x faster (1004 vs 181 docs/s), uses ~3x less VRAM (2.3 vs 7.4 GiB), and halves index size (7.6 vs 15.3 GiB per 1M vectors at native dims).

Why it makes sense

Caveats

Deployed config

1B-AWQ (groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16) as the MemPalace encoder: native 2048D, query: /passage: prefixes, L2-normalized cosine retrieval. Revisit only if real-traffic eval shows Hit@1 gaps the 8B demonstrably closes.