Back to the experimentSupporting notebook

Codec round-trip shootout incl. a new FP8-E4M3 codec (2026-09-21)

Source: moss-tts-awq/codec-roundtrip-fp8-2026-09-21/README.md · revision 6550ead3945b

Five real reference clips (argentina-female-2121) pushed through every audio codec variant on disk — encode → decode → measure (waveform SNR + WavLM x-vector cosine vs the original) → side-by-side listening on the matrix site.

Variants: dense FP32 (control) · INT8-SAFE (HQ4-I8-SAFE) · NF4 blockwise · FP8-E4M3 blockwise (built in this experiment) — each as full round-trip and as “FP32 encodes → X decodes” (the production path).

Headline

Contents

Reproduce

python convert_codec_fp8.py   # builds weights/MOSS-Audio-Tokenizer-FP8 (1.76 GiB)
CUDA_VISIBLE_DEVICES=2 python roundtrip_codecs.py   # -> codecdemo/ + results json

Codec checkpoints are workstation-local ($WORKSTATION/MOSS-TTS/weights/MOSS-Audio-Tokenizer*, INT8 artifact in $WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1).

Boundaries

5 clips, one speaker. SNR/cosine are waveform/speaker proxies, not perceptual codec scores (MUSHRA-style listening would be next). Using the FP8 encoder for clone references is untested — audio parity does not imply identical tokens, and conditioning sensitivity is documented. Anti-overwrite guards on all outputs.