Back to the experimentSupporting notebook

moss-tts-hq4 / phase0-fidelity-envelope-2026-09

Source: moss-tts-hq4/phase0-fidelity-envelope-2026-09/README.md · revision 6550ead3945b

Fast HQ4 build (128 samples / 100 iters, symmetric W4A16) under test: greedy text-channel parity 96.5%, 12 GB envelope 3/3 PASS at 6.8 GiB peak, sampled A/B 33/33 OK both models.

See JOURNEY.md for background. Numbers under docs/, audio under samples/.

Results

Test Result
Greedy parity BF16↔HQ4, text channel (10 prompts) 96.5% (7/10 bit-identical over full 1200 tokens)
Greedy parity, audio channels 21% (expected greedy cascade after first argmax flip)
12 GB envelope, short (4.6s audio) PASS, peak 6.65 GiB, RTF 1.88
12 GB envelope, medium (15.1s audio) PASS, peak 6.83 GiB, RTF 1.92
12 GB envelope, long ES (29.6s audio) PASS, peak 6.83 GiB, RTF 0.94
Sampled A/B BF16 33/33 OK
Sampled A/B HQ4-fast 33/33 OK, median duration delta vs BF16: 1.3s

Honest caveat: RTF ~0.9–1.9 was measured on an RTX 3090. The 3060 has ~2.6× less memory bandwidth, so projected 3060 RTF is ~2–2.5 (fits 12 GB, batch/offline use — not realtime). A faster kernel backend (marlin via gptqmodel) is the obvious next lever.

Artifacts here

Path What
docs/parity_fast.json Per-prompt greedy stats (lengths, text/audio match, first-divergence)
docs/parity_fast_seqs.pt Raw greedy (start, ids) tensors, ref + cand (reusable)
docs/envelope.json 12 GB runs: VRAM phases, RTF, tokens
docs/ab_summary.json 33-prompt sampled A/B per-prompt durations + top duration deltas
samples/parity_*.wav Greedy pairs: acronym ref/cand, en-long-1 cand
samples/envelope_*.wav 12 GB short/medium/long outputs
samples/es-conv-1_{bf16,hq4}.wav Sampled Rioplatense pair for blind listening
samples/prompts.txt Matching prompts