Took the FP8 codec validated in codec-roundtrip-fp8-2026-09-21 and deployed it in the live production daemon on the .54 RTX 3060: the artifact dequantizes to BF16 at load, the encoder half is dropped (decode-only, ~1.6 GiB), and the decoder stays resident on cuda:0 next to the warm generator (6.37 GiB). Reference encoding stays on the FP32/CPU codec (the validated clone fix; the FP8 encoder remains untested for conditioning).
Production code: groxaxo/moss54 commit 40ac7f0
(ops/moss54d/moss54d.py, flag --codec-fp8).
Measured (A/B, seed 1234, same reference, .54 RTX 3060 12 GB)
| Metric | Before (CPU fp32 decode) | After (FP8 GPU decode) |
|---|---|---|
| decode_time_s (clone) | 3.36 | 0.24 (14×) |
| decode_time_s (default voice) | 2.76 | 0.23 |
| gen_time_s | 22.9 | 22.91–22.94 (unchanged) |
| wall_s per render (clone) | 35.1 | 24.4 |
| new_tokens / duration_s | 193 / 12.72 | identical |
| audio_health flags | all false | all false |
| daemon VRAM steady | 6.77→8.01 GiB torch-reserved (net +1.6) |
Verification per the golden protocol: identical sample counts and token/duration parity (PCM mean-abs diff 0.0018 — codec-precision level); ASR A/B (whisperx medium es) identical transcripts before/after in both voices (0.0625/0.0625, 0.0938/0.0938 — residual = transcriber mishears present in both). First post-restart render shows warmup-inflated gen_time; repeat runs are the reportable number.
Contents
docs/before_results.json,docs/after_results.json— raw A/B metrics.docs/wer_ab.json— whisperx transcripts + WER, all four files.samples/— the four A/B WAVs.scripts/render_ab.py— the exact harness (POST /v1/audio/speech).
Boundaries
Single seed, one reference, one text per voice; decode speedup scales with audio length. RTF gen unchanged (7–8 tok/s). The FP8 encoder path is deliberately NOT exercised. No blind ABX listening — parity rests on the round-trip shootout + identical ASR transcripts + unchanged speaker path.