Journal

Round-trip tests expose a broken encoder

The new 1.76GB FP8-E4M3 codec reached FP32-control parity in the measured round-trip test with its encoder intact.

Vlad / experimentos.
The boundary

INT8-SAFE round-trip SNR was about zero. Conclusions apply to the named clips and implementations.

The question

A round-trip test challenged the assumption that a codec was healthy because its decoder worked. The record compares the encoder-plus-decoder path and qualifies the new FP8 implementation on the named clips.

From the original notebook

Five real reference clips (argentina-female-2121) pushed through every audio codec variant on disk — encode → decode → measure (waveform SNR + WavLM x-vector cosine vs the original) → side-by-side listening on the matrix site.

Variants: dense FP32 (control) · INT8-SAFE (HQ4-I8-SAFE) · NF4 blockwise · FP8-E4M3 blockwise (built in this experiment) — each as full round-trip and as “FP32 encodes → X decodes” (the production path).

Headline

  • New FP8-E4M3 codec reaches parity with the FP32 control (SNR 10.7–12.5 dB, cos 0.951–0.994) at 1.76 GB with a working encoder — built with the same RTN blockwise recipe as NF4 (block=64, fp32 absmax; worst weight err 3.2e-2 vs NF4’s 0.119). “NF8” is not a real format; E4M3 is the standard 8-bit float.
  • The INT8-SAFE encoder is broken, now with a number: full round-trip SNR ≈ 0 dB (cos 0.56–0.61) while its decoder is sample-parity with FP32 — the measured root cause of the earlier clone failure (FP32 reference encoding is mandatory).
  • NF4 (0.96 GB) is the only sub-GB codec and works end-to-end, at ~1.5 dB cost vs control.

Contents

Reproduce

python convert_codec_fp8.py   # builds weights/MOSS-Audio-Tokenizer-FP8 (1.76 GiB)
CUDA_VISIBLE_DEVICES=2 python roundtrip_codecs.py   # -> codecdemo/ + results json

Codec checkpoints are workstation-local ($WORKSTATION/MOSS-TTS/weights/MOSS-Audio-Tokenizer*, INT8 artifact in $WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1).

Boundaries

5 clips, one speaker. SNR/cosine are waveform/speaker proxies, not perceptual codec scores (MUSHRA-style listening would be next). Using the FP8 encoder for clone references is untested — audio parity does not imply identical tokens, and conditioning sensitivity is documented. Anti-overwrite guards on all outputs.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS