Back to the experimentSupporting notebook

Findings — voice-cloning failure and the FP32 reference-encoder fix

Source: moss-tts-awq/clone-fp32-reference-2026-09-21/docs/FINDINGS.md · revision 6550ead3945b

Scope: MOSS-TTS-v1.5 AWQ INT4 + INT8 codec (HQ4-I8-SAFE) on RTX 3090, zero-shot cloning of argentina-female-2121 (raw dataset WAVs under $WORKSTATION/train/datasets/latam-es/argentina/), Spanish narration. All runs use the frozen conda env moss-awq3060-w8a8, seed 1234, MOSS-default sampling.

1. Timeline of the failure

Run Reference handling Result
1 (native, no ref) none OK; warm pooled RTF(gen) 0.326 (stories_int8codec.json)
2 (single-shot clone) 12.93 s joined ref, INT8-encoded User heard a male voice → clone failed (stories_clone2121.json)
3 (per-sentence clone) same ref re-injected per sentence, INT8-encoded Still failed; similarity decays 0.955→0.675 mid-story (stories_clone2121_perchunk.json, speaker_scores.json)

Two harness bugs were fixed along the way (kept for the record): messages[0].audio_codes_list[0] discarded any extra decoded segments, and the processor’s batch decode hit the codec’s chunk_duration batch>1 restriction. Neither was the root cause — all observed outputs had a single segment.

2. Root cause: the INT8 codec also encoded the reference

Both story harnesses swapped processor.audio_tokenizer to the INT8 codec before encoding the reference, so quantization affected conditioning, not just output decoding.

Measured on the real path (single-file encode, audit_clone_logic_v4):

3. Controlled isolation (generator × reference encoder × decoder)

compare_clone_precision.py, same two-sentence text, seed 1234; each token stream decoded with both codecs. Speaker metric: WavLM-base-plus-sv x-vector cosine vs the original champú clip (early/mid/late 8 s windows).

Generator Ref encoder Decoder Target cosine
AWQ INT4 INT8 INT8 0.8475
AWQ INT4 FP32 INT8 0.9556
BF16 INT8 INT8 0.8222
BF16 FP32 INT8 0.9591
AWQ INT4 FP32 FP32 0.9651 (champú ref)
BF16 FP32 FP32 0.9665 (champú ref)

Context anchors: two raw clips of the same speaker 0.9674; different female 0.8871; male control 0.7498. The broken short sample scored 0.7078 to target and 0.8869 to the male control — objectively the wrong voice.

Conclusions: reference encoder dominates; output decoder (INT8 vs FP32) is ≈0.001; AWQ vs BF16 generator is small once conditioning is clean.

4. Fix and corrected full stories

Fix: encode the reference once in FP32 before installing the INT8 codec, cache codes on CPU, decode outputs with INT8. Default in both harnesses; --reference-encoder int8 kept only to reproduce the failure.

Regenerated stories (stories_clone2121_fp32ref.json), 124.16 s total, warm pooled RTF(gen) 0.3665:

Story Audio RTF(gen) Similarity, full coverage (8 s/4 s hop)
El pichón y la tormenta 41.36 s 0.4288 0.9587–0.9682 (10 windows)
La abuela y el mate 33.68 s 0.3337 0.9376–0.9575 (8 windows)
El perro que contaba estrellas 49.12 s 0.3365 0.9441–0.9693 (12 windows)

Every window is closer to the target than to both control speakers. Listening acceptance is separate from these scores.

5. Accent: still open

The user then reported Spain-like pronunciation. Verified prompt for all delivered runs: language="Spanish", instruction=None, no system turn, user turn only. Nothing requested Latin American/Argentine pronunciation; the accent was implicitly expected to come from the reference. WavLM cosine does not measure accent. Proposed but untested next step: keep language="Spanish" and pass an explicit instruction (Argentine/LatAm, seseo, preserve reference accent), then evaluate by listening.

6. Reference length: 3.1 s → 12.5 s → 30.4 s

Same-speaker references built from complete clips (no cuts, no pitch/time/loudness changes, 120 ms gaps), speaker verified by full dataset ID plus embeddings:

Matched A/B (same text/seed/prompt, FP32 encoder): 3.1 s → 0.9617; 12.5 s → 0.9407 (reference_3vs12_results.json); 12.5 s → 0.9588 vs 30.4 s → 0.9486 (reference_12vs30_results.json). No similarity gain from longer references in these single pairs — user preferred 12.5 s by ear. Duration alone is not the accent lever.

7. Limits

Evidence index

File What
results.json Factorial generator/encoder/decoder matrix + hashes
speaker_scores.json WavLM evaluation of broken outputs, controls, round-trips
speaker_scores_fp32ref*.json Corrected stories, 3-window + full coverage
stories_*.json Run metrics for runs 1–3 and corrected run
reference_3vs12_results.json, reference_12vs30_results.json Reference-length A/Bs
ref_2121_verified_*.json, ref_2121_30s_candidates.json Reference construction evidence