Back to the experimentSupporting notebook

MOSS-TTS-v1.5 HQ4 — Codec A/B/C: FP32 vs BF16 vs INT8 (2026-09)

Source: moss-tts-hq4/codec-abc-2026-09/JOURNEY.md · revision 6550ead3945b

Background

After the HQ4 fast build (INT4 LLM backbone + FP32 codec), the codec was the biggest remaining chunk: 1.77B params, 6.6 GB FP32. Two reductions were tested against the FP32 reference on identical generations (seed 1234, same sampling):

Test GPUs were blocked (EXL3 server on all 3x3090); user authorized stopping it. The proxy (qwen36_variant_proxy, pid 4128370) was SIGSTOP-frozen during tests to prevent backend relaunch, then SIGCONT-resumed (self-heals on next request — its normal rotation behavior).

Journey

  1. BF16 forward crashed: Input type (float) and bias type (bfloat16). Root cause: the modeling file hard-casts the quantizer core to FP32 (.float() everywhere — exact nearest-neighbor by design), but the BF16 checkpoint has BF16 in/out/output projections. No reduced-precision codec could EVER run without either FP32 quantizer convs or modeling casts.
  2. Fix: 9-site _match_proj_dtype patch in the vendored modeling_moss_audio_tokenizer.py (conv input → weight dtype; encoder-entry cast for the FP32 waveform). Applied identically to both codec dirs. Proven no-op for FP32: fp32 output bit-identical before/after.
  3. Pipeline proven bit-deterministic (two fp32 runs, separate processes → identical bytes). So pure-TTS legs share token sequences exactly; measured diffs are 100% codec decode drift. Clean science.
  4. clone_en confounded: each leg encodes the voice reference with its own codec → different ref codes → different takes (lengths differ). A fair clone comparison needs shared ref codes (harness work, deferred).
  5. INT8: converter (tools/convert_codec_int8.py, worst weight err 8.9e-3)
    • reference loader (tools/codec_int8_loader.py, dequantizes to BF16 compute)
    • --codec-int8 harness flag. INT8 adds ~40-60% more mean drift than BF16 (same order of magnitude). RTF unchanged. NOTE: the reference loader materializes BF16 on GPU, so runtime VRAM == BF16 leg; the 4x win is disk (+ load time). A real INT8 kernel (bnb/marlin-class) is the follow-up for runtime VRAM.
  6. Deliveries: 12 WAVs to Hermes Telegram (batch 1: msgs 11806-11814, batch 2: 11824-11828; 11816-11823 were accidental script re-sends).

Numbers (benchmarks/codec_ab, codec_abc.json)

Pure-TTS legs, same tokens (max/mean abs sample diff):

Maxima are isolated spikes; means ~5e-4-1.7e-3. The 5e-3 max-diff gate FAILs for both — unrealistic bar for a 1.7B BF16 decoder; ears are the arbiter.

WhisperX verdict + INT8 adoption

WhisperX (large-v3-turbo, float16, wav2vec2 alignment, per-case language) transcribed all 12 WAVs. Normalized WER int8-vs-fp32 = 0.000 on all 4 cases (0/16, 0/39, 0/66, 0/13 words — even the uncontrolled clone take speaks identical words). bf16-vs-fp32 also 0.000. The only vs-prompt deviation is long_es 0.031 (2/65), shared identically by all three legs → codec-independent (TTS/ASR artifact, not a quant effect).

Decision: INT8 is content-perfect → adopted as default. infer_hybrid_3060.py now auto-detects INT8 dirs (manifest marker, placeholder config donor) and defaults --codec to weights/MOSS-Audio-Tokenizer-INT8. Default-path output verified bit-identical to the validated short_int8.wav. Bench/test tools keep explicit FP32 defaults (reference role).

Open