Scope: MOSS-TTS-v1.5 AWQ INT4 + INT8 codec (HQ4-I8-SAFE) on RTX 3090,
zero-shot cloning of argentina-female-2121 (raw dataset WAVs under
$WORKSTATION/train/datasets/latam-es/argentina/), Spanish narration. All runs use the
frozen conda env moss-awq3060-w8a8, seed 1234, MOSS-default sampling.
1. Timeline of the failure
| Run | Reference handling | Result |
|---|---|---|
| 1 (native, no ref) | none | OK; warm pooled RTF(gen) 0.326 (stories_int8codec.json) |
| 2 (single-shot clone) | 12.93 s joined ref, INT8-encoded | User heard a male voice → clone failed (stories_clone2121.json) |
| 3 (per-sentence clone) | same ref re-injected per sentence, INT8-encoded | Still failed; similarity decays 0.955→0.675 mid-story (stories_clone2121_perchunk.json, speaker_scores.json) |
Two harness bugs were fixed along the way (kept for the record):
messages[0].audio_codes_list[0] discarded any extra decoded segments, and the
processor’s batch decode hit the codec’s chunk_duration batch>1 restriction.
Neither was the root cause — all observed outputs had a single segment.
2. Root cause: the INT8 codec also encoded the reference
Both story harnesses swapped processor.audio_tokenizer to the INT8 codec
before encoding the reference, so quantization affected conditioning, not
just output decoding.
Measured on the real path (single-file encode, audit_clone_logic_v4):
- FP32 vs INT8 discrete token agreement on the joined reference: 7.7 % (first codebook 7.5 %). Not a similarity score, but the conditioning tokens are almost entirely different.
- INT8 reference round-trip target similarity: 0.565 (joined) / 0.577 (champú) vs FP32 round-trips 0.966 / 0.990.
- Codec codebooks/lookup are bit-identical between the two codecs (98/98 tensors), so this is not a codebook-numbering artifact — it is the INT8 encoder path.
3. Controlled isolation (generator × reference encoder × decoder)
compare_clone_precision.py, same two-sentence text, seed 1234; each token
stream decoded with both codecs. Speaker metric: WavLM-base-plus-sv x-vector
cosine vs the original champú clip (early/mid/late 8 s windows).
| Generator | Ref encoder | Decoder | Target cosine |
|---|---|---|---|
| AWQ INT4 | INT8 | INT8 | 0.8475 |
| AWQ INT4 | FP32 | INT8 | 0.9556 |
| BF16 | INT8 | INT8 | 0.8222 |
| BF16 | FP32 | INT8 | 0.9591 |
| AWQ INT4 | FP32 | FP32 | 0.9651 (champú ref) |
| BF16 | FP32 | FP32 | 0.9665 (champú ref) |
Context anchors: two raw clips of the same speaker 0.9674; different female 0.8871; male control 0.7498. The broken short sample scored 0.7078 to target and 0.8869 to the male control — objectively the wrong voice.
Conclusions: reference encoder dominates; output decoder (INT8 vs FP32) is ≈0.001; AWQ vs BF16 generator is small once conditioning is clean.
4. Fix and corrected full stories
Fix: encode the reference once in FP32 before installing the INT8 codec,
cache codes on CPU, decode outputs with INT8. Default in both harnesses;
--reference-encoder int8 kept only to reproduce the failure.
Regenerated stories (stories_clone2121_fp32ref.json), 124.16 s total, warm
pooled RTF(gen) 0.3665:
| Story | Audio | RTF(gen) | Similarity, full coverage (8 s/4 s hop) |
|---|---|---|---|
| El pichón y la tormenta | 41.36 s | 0.4288 | 0.9587–0.9682 (10 windows) |
| La abuela y el mate | 33.68 s | 0.3337 | 0.9376–0.9575 (8 windows) |
| El perro que contaba estrellas | 49.12 s | 0.3365 | 0.9441–0.9693 (12 windows) |
Every window is closer to the target than to both control speakers. Listening acceptance is separate from these scores.
5. Accent: still open
The user then reported Spain-like pronunciation. Verified prompt for all
delivered runs: language="Spanish", instruction=None, no system turn, user
turn only. Nothing requested Latin American/Argentine pronunciation; the accent
was implicitly expected to come from the reference. WavLM cosine does not
measure accent. Proposed but untested next step: keep language="Spanish"
and pass an explicit instruction (Argentine/LatAm, seseo, preserve reference
accent), then evaluate by listening.
6. Reference length: 3.1 s → 12.5 s → 30.4 s
Same-speaker references built from complete clips (no cuts, no pitch/time/loudness changes, 120 ms gaps), speaker verified by full dataset ID plus embeddings:
ref_2121_verified_12s_v2.wav— 12.532 s, 3 clips, min pairwise 0.957.ref_2121_verified_30s.wav— 30.404 s, 7 clips (12.5 s prefix bit-identical), min pairwise 0.957. 40 candidates scored, 20 eligible, best-combination selection; evidence inref_2121_30s_candidates.json.
Matched A/B (same text/seed/prompt, FP32 encoder):
3.1 s → 0.9617; 12.5 s → 0.9407 (reference_3vs12_results.json);
12.5 s → 0.9588 vs 30.4 s → 0.9486 (reference_12vs30_results.json).
No similarity gain from longer references in these single pairs — user
preferred 12.5 s by ear. Duration alone is not the accent lever.
7. Limits
- WavLM cosine: speaker resemblance proxy; not a gender classifier, accent score, or calibrated identity threshold. Controls are local, not universal.
- Single-pair A/Bs, single seed; Marlin/bf16 is run-to-run nondeterministic.
- The fix bypasses the INT8 encoder path; offending layers were not isolated.
- Reference provenance: raw dataset WAVs, not Breeze-generated outputs from
$WORKSTATION/train/breeze-tts-samples-mf/. Bare2121also exists as a Chilean male in another manifest — always keep country/gender/speaker/file identifiers.
Evidence index
| File | What |
|---|---|
results.json |
Factorial generator/encoder/decoder matrix + hashes |
speaker_scores.json |
WavLM evaluation of broken outputs, controls, round-trips |
speaker_scores_fp32ref*.json |
Corrected stories, 3-window + full coverage |
stories_*.json |
Run metrics for runs 1–3 and corrected run |
reference_3vs12_results.json, reference_12vs30_results.json |
Reference-length A/Bs |
ref_2121_verified_*.json, ref_2121_30s_candidates.json |
Reference construction evidence |