Fast HQ4 build (128 samples / 100 iters, symmetric W4A16) under test: greedy text-channel parity 96.5%, 12 GB envelope 3/3 PASS at 6.8 GiB peak, sampled A/B 33/33 OK both models.
See JOURNEY.md for background. Numbers under docs/, audio under samples/.
Results
| Test | Result |
|---|---|
| Greedy parity BF16↔HQ4, text channel (10 prompts) | 96.5% (7/10 bit-identical over full 1200 tokens) |
| Greedy parity, audio channels | 21% (expected greedy cascade after first argmax flip) |
| 12 GB envelope, short (4.6s audio) | PASS, peak 6.65 GiB, RTF 1.88 |
| 12 GB envelope, medium (15.1s audio) | PASS, peak 6.83 GiB, RTF 1.92 |
| 12 GB envelope, long ES (29.6s audio) | PASS, peak 6.83 GiB, RTF 0.94 |
| Sampled A/B BF16 | 33/33 OK |
| Sampled A/B HQ4-fast | 33/33 OK, median duration delta vs BF16: 1.3s |
Honest caveat: RTF ~0.9–1.9 was measured on an RTX 3090. The 3060 has ~2.6× less memory bandwidth, so projected 3060 RTF is ~2–2.5 (fits 12 GB, batch/offline use — not realtime). A faster kernel backend (marlin via gptqmodel) is the obvious next lever.
Artifacts here
| Path | What |
|---|---|
docs/parity_fast.json |
Per-prompt greedy stats (lengths, text/audio match, first-divergence) |
docs/parity_fast_seqs.pt |
Raw greedy (start, ids) tensors, ref + cand (reusable) |
docs/envelope.json |
12 GB runs: VRAM phases, RTF, tokens |
docs/ab_summary.json |
33-prompt sampled A/B per-prompt durations + top duration deltas |
samples/parity_*.wav |
Greedy pairs: acronym ref/cand, en-long-1 cand |
samples/envelope_*.wav |
12 GB short/medium/long outputs |
samples/es-conv-1_{bf16,hq4}.wav |
Sampled Rioplatense pair for blind listening |
samples/prompts.txt |
Matching prompts |