Back to the experimentSupporting notebook

Quantization matrix for cloned Spanish stories (2026-09-21)

Source: moss-tts-awq/quant-matrix-stories-2026-09-21/README.md · revision 6550ead3945b

The three benchmark stories, regenerated under every generator quantization × output codec available on the workstation, from one frozen setup: same text, seed 1234, MOSS-default sampling, user-approved 12.532 s reference (argentina-female-2121), FP32 reference encoding (the validated fix).

Axis Variants
Generator BF16 dense · HQ4 (AutoRound INT4) · AWQ INT4 (Marlin)
Output codec dense FP32 · HQ4-I8-SAFE INT8

Prompt unchanged across cells: language="Spanish", instruction=None (no accent instruction — the open Spain-accent issue applies to all cells).

Headline

Contents

Reproduce

CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  WORKSTATION/python run_stories_quant_matrix.py \
  --generator bf16 --decoder fp32   # {bf16,hq4,awq} × {fp32,int8}

CUDA_VISIBLE_DEVICES=1 WORKSTATION/python score_quant_matrix.py
WORKSTATION/python build_quant_site.py

Models are workstation-local ($WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5, $WORKSTATION/models/MOSS-TTS-v1.5-{AWQ4,HQ4}, INT8 codec artifact in $WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1). Anti-overwrite guards: output dirs and metrics files refuse to replace existing runs.

Boundaries

Single seed, single reference, one story set; Marlin/bf16 nondeterminism applies. WavLM cosine is a local speaker-resemblance diagnostic — not accent, gender, or calibrated identity. VRAM peaks are gen-phase peaks with this exact offload configuration. Listening verdicts live with the user.