The three benchmark stories, regenerated under every generator quantization ×
output codec available on the workstation, from one frozen setup: same text,
seed 1234, MOSS-default sampling, user-approved 12.532 s reference
(argentina-female-2121), FP32 reference encoding (the validated fix).
| Axis | Variants |
|---|---|
| Generator | BF16 dense · HQ4 (AutoRound INT4) · AWQ INT4 (Marlin) |
| Output codec | dense FP32 · HQ4-I8-SAFE INT8 |
Prompt unchanged across cells: language="Spanish", instruction=None
(no accent instruction — the open Spain-accent issue applies to all cells).
Headline
- Speaker similarity is effectively codec-independent (FP32 vs INT8 decoder ≤0.001 difference per story, replicating the earlier factorial), and ranks generators BF16 ≥ AWQ ≳ HQ4 (means 0.958–0.965 / 0.949–0.960 / 0.953–0.960; anchors: same-speaker 0.9674, other-female 0.8871, male 0.7498). HQ4 logged the lowest single window (0.915). None of this measures accent or naturalness.
- Pacing differs audibly: identical text yields 131.7 s (HQ4), 117.8 s (BF16), 110–112 s (AWQ) total — the quantizations speak at different rates.
- Cost: AWQ is fastest (pooled RTF(gen) 0.332–0.338, 6.6 GiB); HQ4 0.355–0.358; BF16 0.396–0.401 at 16.2 GiB (fp32 codec) — and BF16+FP32-codec only fits via CPU-offloading the codec between decodes (first attempt OOMed at model load).
- Listening site (static, no JS/telemetry) served at
private workstation servicewith all 18 renderings + real references side by side. - Listener verdict (single user, same day): AWQ INT4 “excellent, basically no difference” vs BF16 dense — subjective, not blind ABX, but combined with the 2.5x smaller footprint it settles AWQ INT4 as the deployment default for this voice. Accent (LatAm vs peninsular) received no explicit verdict.
Contents
- docs/QUANT-MATRIX-REVIEW.md — full tables, OOM incident + fix, site details.
docs/<cell>.json×6 — per-cell run metrics (RTF, VRAM, tokens, hashes).docs/speaker_scores.json— WavLM full-coverage evaluation, 18 outputs.scripts/— one-cell-per-run harness, matrix scorer, static site builder.site-audio-sample/— renderedindex.html(audio paths are workstation-local) plus one reference and two story-1 MP3s (BF16 vs AWQ, INT8 codec) as frozen examples.
Reproduce
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
WORKSTATION/python run_stories_quant_matrix.py \
--generator bf16 --decoder fp32 # {bf16,hq4,awq} × {fp32,int8}
CUDA_VISIBLE_DEVICES=1 WORKSTATION/python score_quant_matrix.py
WORKSTATION/python build_quant_site.py
Models are workstation-local ($WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5,
$WORKSTATION/models/MOSS-TTS-v1.5-{AWQ4,HQ4}, INT8 codec artifact in
$WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1). Anti-overwrite
guards: output dirs and metrics files refuse to replace existing runs.
Boundaries
Single seed, single reference, one story set; Marlin/bf16 nondeterminism applies. WavLM cosine is a local speaker-resemblance diagnostic — not accent, gender, or calibrated identity. VRAM peaks are gen-phase peaks with this exact offload configuration. Listening verdicts live with the user.