Background
After the HQ4 fast build (INT4 LLM backbone + FP32 codec), the codec was the biggest remaining chunk: 1.77B params, 6.6 GB FP32. Two reductions were tested against the FP32 reference on identical generations (seed 1234, same sampling):
- BF16:
weights/MOSS-Audio-Tokenizer-BF16(3.4 GB, 32 codebooks kept FP32) - INT8:
weights/MOSS-Audio-Tokenizer-INT8(1.66 GB, symmetric per-channel RTN over all 554 Linears + 66 weight-norm Conv1ds = 99.9% of params, codebooks FP32, biases/norms BF16). No calibration (RTN needs none).
Test GPUs were blocked (EXL3 server on all 3x3090); user authorized stopping it. The proxy (qwen36_variant_proxy, pid 4128370) was SIGSTOP-frozen during tests to prevent backend relaunch, then SIGCONT-resumed (self-heals on next request — its normal rotation behavior).
Journey
- BF16 forward crashed:
Input type (float) and bias type (bfloat16). Root cause: the modeling file hard-casts the quantizer core to FP32 (.float()everywhere — exact nearest-neighbor by design), but the BF16 checkpoint has BF16 in/out/output projections. No reduced-precision codec could EVER run without either FP32 quantizer convs or modeling casts. - Fix: 9-site
_match_proj_dtypepatch in the vendoredmodeling_moss_audio_tokenizer.py(conv input → weight dtype; encoder-entry cast for the FP32 waveform). Applied identically to both codec dirs. Proven no-op for FP32: fp32 output bit-identical before/after. - Pipeline proven bit-deterministic (two fp32 runs, separate processes → identical bytes). So pure-TTS legs share token sequences exactly; measured diffs are 100% codec decode drift. Clean science.
- clone_en confounded: each leg encodes the voice reference with its own codec → different ref codes → different takes (lengths differ). A fair clone comparison needs shared ref codes (harness work, deferred).
- INT8: converter (
tools/convert_codec_int8.py, worst weight err 8.9e-3)- reference loader (
tools/codec_int8_loader.py, dequantizes to BF16 compute) --codec-int8harness flag. INT8 adds ~40-60% more mean drift than BF16 (same order of magnitude). RTF unchanged. NOTE: the reference loader materializes BF16 on GPU, so runtime VRAM == BF16 leg; the 4x win is disk (+ load time). A real INT8 kernel (bnb/marlin-class) is the follow-up for runtime VRAM.
- reference loader (
- Deliveries: 12 WAVs to Hermes Telegram (batch 1: msgs 11806-11814, batch 2: 11824-11828; 11816-11823 were accidental script re-sends).
Numbers (benchmarks/codec_ab, codec_abc.json)
Pure-TTS legs, same tokens (max/mean abs sample diff):
- short: bf16/fp32 0.133/1.24e-3, int8/fp32 0.214/1.69e-3, int8/bf16 0.142/1.45e-3
- medium: bf16/fp32 0.092/5.05e-4, int8/fp32 0.109/7.39e-4, int8/bf16 0.068/5.08e-4
- long_es: bf16/fp32 0.034/5.52e-4, int8/fp32 0.040/8.06e-4, int8/bf16 0.024/7.25e-4
Maxima are isolated spikes; means ~5e-4-1.7e-3. The 5e-3 max-diff gate FAILs for both — unrealistic bar for a 1.7B BF16 decoder; ears are the arbiter.
WhisperX verdict + INT8 adoption
WhisperX (large-v3-turbo, float16, wav2vec2 alignment, per-case language) transcribed all 12 WAVs. Normalized WER int8-vs-fp32 = 0.000 on all 4 cases (0/16, 0/39, 0/66, 0/13 words — even the uncontrolled clone take speaks identical words). bf16-vs-fp32 also 0.000. The only vs-prompt deviation is long_es 0.031 (2/65), shared identically by all three legs → codec-independent (TTS/ASR artifact, not a quant effect).
Decision: INT8 is content-perfect → adopted as default. infer_hybrid_3060.py
now auto-detects INT8 dirs (manifest marker, placeholder config donor) and
defaults --codec to weights/MOSS-Audio-Tokenizer-INT8. Default-path output
verified bit-identical to the validated short_int8.wav. Bench/test tools keep
explicit FP32 defaults (reference role).
Open
- Blind-listening verdict (user) — WhisperX says perfect, ears confirm.
- Fair clone A/B with shared ref codes (optional harness work).
- INT8 runtime kernels for real VRAM win (optional).
- Full-spec HQ4 rebuild + manifest (parked per user).