The question
The clone quality problem was narrowed to the reference encoder. Changing that stage to FP32 recovered the measured target similarity while the output decoder stayed INT8, leaving accent and listening questions open.
From the original notebook
Zero-shot voice cloning with the AWQ INT4 generator and the HQ4-I8-SAFE codec produced the wrong voice (male) from a female Argentine reference. A controlled factorial comparison (generator × reference encoder × output decoder) isolated the cause: the INT8 codec was also encoding the reference audio, corrupting the conditioning tokens. Encoding the reference in FP32 and keeping INT8 only for output decoding restored target-speaker similarity from 0.8475 → 0.9556 (AWQ) and 0.8222 → 0.9591 (BF16 generator).
Working directory with live artifacts: $WORKSTATION/moss-awq4-stories/ (reports:
CLONING-REVIEW.md, CLONING-FIX-FP32-REFERENCE.md, PROMPT-ACCENT-REVIEW.md,
REFERENCE-12S-REVIEW.md, REFERENCE-30S-REVIEW.md, TEST-REPORT.md).
Headline
- Root cause: INT8 reference encoding — only 7.7 % token agreement with the FP32 encoder on the same WAV; INT8 reference round-trip similarity 0.57 vs 0.97–0.99 FP32. Codebooks are bit-identical (98/98), so it is the encoder path, not numbering.
- Fix (validated): FP32-encode the reference once, cache the codes, decode outputs with INT8. Default in both harnesses. Output decoder choice moves similarity by ≈0.001; generator AWQ-vs-BF16 is small once conditioning is clean.
- Corrected stories: 124.16 s across 3 stories, warm pooled RTF(gen) 0.3665; every full-coverage window (8 s / 4 s hop, 30 windows) closer to the target than to control speakers (0.938–0.969).
- Open: Spain-like accent reported by listener. Prompt audit confirmed
language="Spanish"+instruction=None— no accent instruction was ever passed. Explicit LatAm/Argentine instruction is proposed but untested; WavLM cosine does not measure accent. - Reference length: verified same-speaker references at 3.116 s / 12.532 s / 30.404 s; matched A/Bs show no similarity gain from longer references (12.5 s preferred by ear by the user).
Contents
docs/FINDINGS.md— full findings, evidence tables, limits.docs/*.json— factorial control matrix, speaker evaluations (3-window and full coverage), story-run metrics, reference-construction evidence.samples/— broken vs fixed OGG pair (same text/seed, only the reference encoder changed), corrected story 1, the three real reference recordings (3.1 s champú, 12.5 s, 30.4 s), and the 12-vs-30 generated A/B.scripts/— the exact harness used (story runners with cached/verified references, precision factorial, speaker scorer, reference builders, A/B tester).
| Sample | What to listen for |
|---|---|
broken_awq_joinedref_int8enc.ogg vs fixed_awq_joinedref_fp32enc.ogg |
Same text/seed; only reference encoder differs |
fixed_story1_fp32ref.ogg |
Corrected full story (FP32 ref, AWQ gen, INT8 decode) |
reference_original_champu_003449.ogg / ref_2121_verified_12s_v2.ogg / ref_2121_verified_30s.ogg |
Real reference recordings, not generated speech |
A_reference_12s.ogg / B_reference_30s.ogg |
Reference-length A/B, same text/seed/prompt |
Reproduce
## corrected full stories (idle GPU 2, frozen env)
CUDA_VISIBLE_DEVICES=2 PYTORCH_ALLOC_CONF=expandable_segments:True \
WORKSTATION/python run_stories_int8codec.py \
--reference WORKSTATION/argentina-female-2121-003449.wav \
--reference-encoder fp32 --prefix fp32ref_ \
--outdir out_clone2121_fp32ref --metrics logs/stories_clone2121_fp32ref.json
## factorial isolation of generator/encoder/decoder
python compare_clone_precision.py # -> precision_controls/results.json
## speaker-embedding evaluation (8 s windows, 4 s hop)
python score_clone_speakers.py --new-only --full-coverage --output speaker_scores_new.json
Models/paths are workstation-local: $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4,
$WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5 (BF16 control),
$WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1 (INT8 codec),
speaker model microsoft/wavlm-base-plus-sv (HF cache; HF_HUB_DISABLE_XET=1).
Boundaries
Measured-result vs interpretation discipline applies (see repo README). WavLM cosine is a local speaker-resemblance proxy — not an accent, gender, or calibrated identity test; listening acceptance is tracked separately and the accent issue remains open. Single-pair A/Bs, single seed, Marlin/bf16 nondeterminism caveat. The INT8 encoder’s offending layers were not isolated; the fix bypasses that path.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.