The question
Codec precision can change storage, waveform detail and reference encoding in different ways. The A/B/C experiment separated those effects using pure speech and clone cases, rather than letting transcription parity stand in for every kind of fidelity.
From the original notebook
Background
After the HQ4 fast build (INT4 LLM backbone + FP32 codec), the codec was the biggest remaining chunk: 1.77B params, 6.6 GB FP32. Two reductions were tested against the FP32 reference on identical generations (seed 1234, same sampling):
- BF16:
weights/MOSS-Audio-Tokenizer-BF16(3.4 GB, 32 codebooks kept FP32) - INT8:
weights/MOSS-Audio-Tokenizer-INT8(1.66 GB, symmetric per-channel RTN over all 554 Linears + 66 weight-norm Conv1ds = 99.9% of params, codebooks FP32, biases/norms BF16). No calibration (RTN needs none).
Test GPUs were blocked (EXL3 server on all 3x3090); user authorized stopping it. The proxy (qwen36_variant_proxy, pid 4128370) was SIGSTOP-frozen during tests to prevent backend relaunch, then SIGCONT-resumed (self-heals on next request — its normal rotation behavior).
Journey
- BF16 forward crashed:
Input type (float) and bias type (bfloat16). Root cause: the modeling file hard-casts the quantizer core to FP32 (.float()everywhere — exact nearest-neighbor by design), but the BF16 checkpoint has BF16 in/out/output projections. No reduced-precision codec could EVER run without either FP32 quantizer convs or modeling casts. - Fix: 9-site
_match_proj_dtypepatch in the vendoredmodeling_moss_audio_tokenizer.py(conv input → weight dtype; encoder-entry cast for the FP32 waveform). Applied identically to both codec dirs. Proven no-op for FP32: fp32 output bit-identical before/after. - Pipeline proven bit-deterministic (two fp32 runs, separate processes → identical bytes). So pure-TTS legs share token sequences exactly; measured diffs are 100% codec decode drift. Clean science.
- clone_en confounded: each leg encodes the voice reference with its own codec → different ref codes → different takes (lengths differ). A fair clone comparison needs shared ref codes (harness work, deferred).
- INT8: converter (
tools/convert_codec_int8.py, worst weight err 8.9e-3)- reference loader (
tools/codec_int8_loader.py, dequantizes to BF16 compute) --codec-int8harness flag. INT8 adds ~40-60% more mean drift than BF16 (same order of magnitude). RTF unchanged. NOTE: the reference loader materializes BF16 on GPU, so runtime VRAM == BF16 leg; the 4x win is disk (+ load time). A real INT8 kernel (bnb/marlin-class) is the follow-up for runtime VRAM.
- reference loader (
- Deliveries: 12 WAVs to Hermes Telegram (batch 1: msgs 11806-11814, batch 2: 11824-11828; 11816-11823 were accidental script re-sends).
Numbers (benchmarks/codec_ab, codec_abc.json)
Pure-TTS legs, same tokens (max/mean abs sample diff):
- short: bf16/fp32 0.133/1.24e-3, int8/fp32 0.214/1.69e-3, int8/bf16 0.142/1.45e-3
- medium: bf16/fp32 0.092/5.05e-4, int8/fp32 0.109/7.39e-4, int8/bf16 0.068/5.08e-4
- long_es: bf16/fp32 0.034/5.52e-4, int8/fp32 0.040/8.06e-4, int8/bf16 0.024/7.25e-4
Maxima are isolated spikes; means ~5e-4-1.7e-3. The 5e-3 max-diff gate FAILs for both — unrealistic bar for a 1.7B BF16 decoder; ears are the arbiter.
WhisperX verdict + INT8 adoption
WhisperX (large-v3-turbo, float16, wav2vec2 alignment, per-case language) transcribed all 12 WAVs. Normalized WER int8-vs-fp32 = 0.000 on all 4 cases (0/16, 0/39, 0/66, 0/13 words — even the uncontrolled clone take speaks identical words). bf16-vs-fp32 also 0.000. The only vs-prompt deviation is long_es 0.031 (2/65), shared identically by all three legs → codec-independent (TTS/ASR artifact, not a quant effect).
Decision: INT8 is content-perfect → adopted as default. infer_hybrid_3060.py
now auto-detects INT8 dirs (manifest marker, placeholder config donor) and
defaults --codec to weights/MOSS-Audio-Tokenizer-INT8. Default-path output
verified bit-identical to the validated short_int8.wav. Bench/test tools keep
explicit FP32 defaults (reference role).
Open
- Blind-listening verdict (user) — WhisperX says perfect, ears confirm.
- Fair clone A/B with shared ref codes (optional harness work).
- INT8 runtime kernels for real VRAM win (optional).
- Full-spec HQ4 rebuild + manifest (parked per user).
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.