FP32 vs BF16 vs INT8 MOSS audio codec, same-token A/B(/C). Pipeline proven bit-deterministic, so pure-TTS legs differ by decode drift only.
| case | bf16/fp32 mean (max) | int8/fp32 mean (max) | int8/bf16 mean |
|---|---|---|---|
| short | 1.24e-3 (0.133) | 1.69e-3 (0.214) | 1.45e-3 |
| medium | 5.05e-4 (0.092) | 7.39e-4 (0.109) | 5.08e-4 |
| long_es | 5.52e-4 (0.034) | 8.06e-4 (0.040) | 7.25e-4 |
| clone_en | uncontrolled (ref encoded per leg — different takes) |
Sizes: FP32 6.6 GB → BF16 3.4 GB (codebooks FP32) → INT8 1.66 GB (RTN per-channel, 620/620 linears+convs = 99.9% params, codebooks FP32). Co-resident peak: 13.0 → 9.6 GiB (BF16/INT8-ref). RTF unchanged (~0.8-1.9 on 3090).
Required a 9-site mixed-precision patch in the vendored modeling file (FP32 quantizer core vs BF16 convs crashed forward; no-op for FP32, proven bit-identical). INT8 runs via a dequantizing reference loader (BF16 compute) — disk win is real, runtime VRAM win needs INT8 kernels (follow-up).
Verdict: numeric gate (max < 5e-3) FAILs for both — that bar was unrealistic for a 1.7B BF16 decoder. WhisperX judge (large-v3-turbo + alignment): int8-vs-fp32 WER = 0.000 on all 4 cases (see whisperx.json) → INT8 adopted as the default codec (harness auto-detects the INT8 dir; default output proven bit-identical to the validated leg). Listening samples with the user (Telegram msgs 11806-11814 + 11824-11828).
See JOURNEY.md for the full story. Samples in samples/.