Back to the experimentSupporting notebook

FP8-E4M3 codec deployed GPU-resident in moss54d — 14× faster decode (2026-09-22)

Source: moss-tts-awq/fp8-codec-deploy-2026-09-22/README.md · revision 6550ead3945b

Took the FP8 codec validated in codec-roundtrip-fp8-2026-09-21 and deployed it in the live production daemon on the .54 RTX 3060: the artifact dequantizes to BF16 at load, the encoder half is dropped (decode-only, ~1.6 GiB), and the decoder stays resident on cuda:0 next to the warm generator (6.37 GiB). Reference encoding stays on the FP32/CPU codec (the validated clone fix; the FP8 encoder remains untested for conditioning).

Production code: groxaxo/moss54 commit 40ac7f0 (ops/moss54d/moss54d.py, flag --codec-fp8).

Measured (A/B, seed 1234, same reference, .54 RTX 3060 12 GB)

Metric Before (CPU fp32 decode) After (FP8 GPU decode)
decode_time_s (clone) 3.36 0.24 (14×)
decode_time_s (default voice) 2.76 0.23
gen_time_s 22.9 22.91–22.94 (unchanged)
wall_s per render (clone) 35.1 24.4
new_tokens / duration_s 193 / 12.72 identical
audio_health flags all false all false
daemon VRAM steady 6.77→8.01 GiB torch-reserved (net +1.6)

Verification per the golden protocol: identical sample counts and token/duration parity (PCM mean-abs diff 0.0018 — codec-precision level); ASR A/B (whisperx medium es) identical transcripts before/after in both voices (0.0625/0.0625, 0.0938/0.0938 — residual = transcriber mishears present in both). First post-restart render shows warmup-inflated gen_time; repeat runs are the reportable number.

Contents

Boundaries

Single seed, one reference, one text per voice; decode speedup scales with audio length. RTF gen unchanged (7–8 tok/s). The FP8 encoder path is deliberately NOT exercised. No blind ABX listening — parity rests on the round-trip shootout + identical ASR transcripts + unchanged speaker path.