The question
The same text made it possible to compare generator and codec precision separately. The six combinations exposed a small decoder similarity effect alongside a much more audible difference in pacing.
From the original notebook
The three benchmark stories, regenerated under every generator quantization ×
output codec available on the workstation, from one frozen setup: same text,
seed 1234, MOSS-default sampling, user-approved 12.532 s reference
(argentina-female-2121), FP32 reference encoding (the validated fix).
| Axis | Variants |
|---|---|
| Generator | BF16 dense · HQ4 (AutoRound INT4) · AWQ INT4 (Marlin) |
| Output codec | dense FP32 · HQ4-I8-SAFE INT8 |
Prompt unchanged across cells: language="Spanish", instruction=None
(no accent instruction — the open Spain-accent issue applies to all cells).
Headline
- Speaker similarity is effectively codec-independent (FP32 vs INT8 decoder ≤0.001 difference per story, replicating the earlier factorial), and ranks generators BF16 ≥ AWQ ≳ HQ4 (means 0.958–0.965 / 0.949–0.960 / 0.953–0.960; anchors: same-speaker 0.9674, other-female 0.8871, male 0.7498). HQ4 logged the lowest single window (0.915). None of this measures accent or naturalness.
- Pacing differs audibly: identical text yields 131.7 s (HQ4), 117.8 s (BF16), 110–112 s (AWQ) total — the quantizations speak at different rates.
- Cost: AWQ is fastest (pooled RTF(gen) 0.332–0.338, 6.6 GiB); HQ4 0.355–0.358; BF16 0.396–0.401 at 16.2 GiB (fp32 codec) — and BF16+FP32-codec only fits via CPU-offloading the codec between decodes (first attempt OOMed at model load).
- Listening site (static, no JS/telemetry) served at
private workstation servicewith all 18 renderings + real references side by side. - Listener verdict (single user, same day): AWQ INT4 “excellent, basically no difference” vs BF16 dense — subjective, not blind ABX, but combined with the 2.5x smaller footprint it settles AWQ INT4 as the deployment default for this voice. Accent (LatAm vs peninsular) received no explicit verdict.
Contents
- docs/QUANT-MATRIX-REVIEW.md — full tables, OOM incident + fix, site details.
docs/<cell>.json×6 — per-cell run metrics (RTF, VRAM, tokens, hashes).docs/speaker_scores.json— WavLM full-coverage evaluation, 18 outputs.scripts/— one-cell-per-run harness, matrix scorer, static site builder.site-audio-sample/— renderedindex.html(audio paths are workstation-local) plus one reference and two story-1 MP3s (BF16 vs AWQ, INT8 codec) as frozen examples.
Reproduce
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
WORKSTATION/python run_stories_quant_matrix.py \
--generator bf16 --decoder fp32 # {bf16,hq4,awq} × {fp32,int8}
CUDA_VISIBLE_DEVICES=1 WORKSTATION/python score_quant_matrix.py
WORKSTATION/python build_quant_site.py
Models are workstation-local ($WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5,
$WORKSTATION/models/MOSS-TTS-v1.5-{AWQ4,HQ4}, INT8 codec artifact in
$WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1). Anti-overwrite
guards: output dirs and metrics files refuse to replace existing runs.
Boundaries
Single seed, single reference, one story set; Marlin/bf16 nondeterminism applies. WavLM cosine is a local speaker-resemblance diagnostic — not accent, gender, or calibrated identity. VRAM peaks are gen-phase peaks with this exact offload configuration. Listening verdicts live with the user.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.