Zero-shot voice cloning with the AWQ INT4 generator and the HQ4-I8-SAFE codec produced the wrong voice (male) from a female Argentine reference. A controlled factorial comparison (generator × reference encoder × output decoder) isolated the cause: the INT8 codec was also encoding the reference audio, corrupting the conditioning tokens. Encoding the reference in FP32 and keeping INT8 only for output decoding restored target-speaker similarity from 0.8475 → 0.9556 (AWQ) and 0.8222 → 0.9591 (BF16 generator).
Working directory with live artifacts: $WORKSTATION/moss-awq4-stories/ (reports:
CLONING-REVIEW.md, CLONING-FIX-FP32-REFERENCE.md, PROMPT-ACCENT-REVIEW.md,
REFERENCE-12S-REVIEW.md, REFERENCE-30S-REVIEW.md, TEST-REPORT.md).
Headline
- Root cause: INT8 reference encoding — only 7.7 % token agreement with the FP32 encoder on the same WAV; INT8 reference round-trip similarity 0.57 vs 0.97–0.99 FP32. Codebooks are bit-identical (98/98), so it is the encoder path, not numbering.
- Fix (validated): FP32-encode the reference once, cache the codes, decode outputs with INT8. Default in both harnesses. Output decoder choice moves similarity by ≈0.001; generator AWQ-vs-BF16 is small once conditioning is clean.
- Corrected stories: 124.16 s across 3 stories, warm pooled RTF(gen) 0.3665; every full-coverage window (8 s / 4 s hop, 30 windows) closer to the target than to control speakers (0.938–0.969).
- Open: Spain-like accent reported by listener. Prompt audit confirmed
language="Spanish"+instruction=None— no accent instruction was ever passed. Explicit LatAm/Argentine instruction is proposed but untested; WavLM cosine does not measure accent. - Reference length: verified same-speaker references at 3.116 s / 12.532 s / 30.404 s; matched A/Bs show no similarity gain from longer references (12.5 s preferred by ear by the user).
Contents
docs/FINDINGS.md— full findings, evidence tables, limits.docs/*.json— factorial control matrix, speaker evaluations (3-window and full coverage), story-run metrics, reference-construction evidence.samples/— broken vs fixed OGG pair (same text/seed, only the reference encoder changed), corrected story 1, the three real reference recordings (3.1 s champú, 12.5 s, 30.4 s), and the 12-vs-30 generated A/B.scripts/— the exact harness used (story runners with cached/verified references, precision factorial, speaker scorer, reference builders, A/B tester).
| Sample | What to listen for |
|---|---|
broken_awq_joinedref_int8enc.ogg vs fixed_awq_joinedref_fp32enc.ogg |
Same text/seed; only reference encoder differs |
fixed_story1_fp32ref.ogg |
Corrected full story (FP32 ref, AWQ gen, INT8 decode) |
reference_original_champu_003449.ogg / ref_2121_verified_12s_v2.ogg / ref_2121_verified_30s.ogg |
Real reference recordings, not generated speech |
A_reference_12s.ogg / B_reference_30s.ogg |
Reference-length A/B, same text/seed/prompt |
Reproduce
## corrected full stories (idle GPU 2, frozen env)
CUDA_VISIBLE_DEVICES=2 PYTORCH_ALLOC_CONF=expandable_segments:True \
WORKSTATION/python run_stories_int8codec.py \
--reference WORKSTATION/argentina-female-2121-003449.wav \
--reference-encoder fp32 --prefix fp32ref_ \
--outdir out_clone2121_fp32ref --metrics logs/stories_clone2121_fp32ref.json
## factorial isolation of generator/encoder/decoder
python compare_clone_precision.py # -> precision_controls/results.json
## speaker-embedding evaluation (8 s windows, 4 s hop)
python score_clone_speakers.py --new-only --full-coverage --output speaker_scores_new.json
Models/paths are workstation-local: $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4,
$WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5 (BF16 control),
$WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1 (INT8 codec),
speaker model microsoft/wavlm-base-plus-sv (HF cache; HF_HUB_DISABLE_XET=1).
Boundaries
Measured-result vs interpretation discipline applies (see repo README). WavLM cosine is a local speaker-resemblance proxy — not an accent, gender, or calibrated identity test; listening acceptance is tracked separately and the accent issue remains open. Single-pair A/Bs, single seed, Marlin/bf16 nondeterminism caveat. The INT8 encoder’s offending layers were not isolated; the fix bypasses that path.