Journal

The clone bug was in the reference encoder

Using FP32 reference encoding lifted AWQ WavLM target similarity from 0.8475 to 0.9556 while retaining INT8 output decode.

Vlad / experimentos.
The boundary

Accent remained open. WavLM similarity is a measured proxy, not a complete human listening verdict.

Measurements

Recorded result

Changing the reference encoder recovered similarity

AWQ generator · INT8 output decoder · named reference test

Changing the reference encoder recovered similarityAWQ + INT8 reference: 0.8475 WavLM target similarity; AWQ + FP32 reference: 0.9556 WavLM target similarity. A similarity proxy is not a human preference or accent verdict.AWQ + INT8 reference0.8475AWQ + INT8 reference: 0.8475 WavLM target similarityAWQ + FP32 reference0.9556AWQ + FP32 reference: 0.9556 WavLM target similarity0WavLM target similarity
  1. AWQ + INT8 reference0.8475
  2. AWQ + FP32 reference0.9556

WavLM target similarity

A similarity proxy is not a human preference or accent verdict.

View data & source
Changing the reference encoder recovered similarity · WavLM target similarity
ConfigurationValue
AWQ + INT8 reference0.8475
AWQ + FP32 reference0.9556

Origin: reported table. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

The clone quality problem was narrowed to the reference encoder. Changing that stage to FP32 recovered the measured target similarity while the output decoder stayed INT8, leaving accent and listening questions open.

From the original notebook

Zero-shot voice cloning with the AWQ INT4 generator and the HQ4-I8-SAFE codec produced the wrong voice (male) from a female Argentine reference. A controlled factorial comparison (generator × reference encoder × output decoder) isolated the cause: the INT8 codec was also encoding the reference audio, corrupting the conditioning tokens. Encoding the reference in FP32 and keeping INT8 only for output decoding restored target-speaker similarity from 0.8475 → 0.9556 (AWQ) and 0.8222 → 0.9591 (BF16 generator).

Working directory with live artifacts: $WORKSTATION/moss-awq4-stories/ (reports: CLONING-REVIEW.md, CLONING-FIX-FP32-REFERENCE.md, PROMPT-ACCENT-REVIEW.md, REFERENCE-12S-REVIEW.md, REFERENCE-30S-REVIEW.md, TEST-REPORT.md).

Headline

  • Root cause: INT8 reference encoding — only 7.7 % token agreement with the FP32 encoder on the same WAV; INT8 reference round-trip similarity 0.57 vs 0.97–0.99 FP32. Codebooks are bit-identical (98/98), so it is the encoder path, not numbering.
  • Fix (validated): FP32-encode the reference once, cache the codes, decode outputs with INT8. Default in both harnesses. Output decoder choice moves similarity by ≈0.001; generator AWQ-vs-BF16 is small once conditioning is clean.
  • Corrected stories: 124.16 s across 3 stories, warm pooled RTF(gen) 0.3665; every full-coverage window (8 s / 4 s hop, 30 windows) closer to the target than to control speakers (0.938–0.969).
  • Open: Spain-like accent reported by listener. Prompt audit confirmed language="Spanish" + instruction=None — no accent instruction was ever passed. Explicit LatAm/Argentine instruction is proposed but untested; WavLM cosine does not measure accent.
  • Reference length: verified same-speaker references at 3.116 s / 12.532 s / 30.404 s; matched A/Bs show no similarity gain from longer references (12.5 s preferred by ear by the user).

Contents

  • docs/FINDINGS.md — full findings, evidence tables, limits.
  • docs/*.json — factorial control matrix, speaker evaluations (3-window and full coverage), story-run metrics, reference-construction evidence.
  • samples/ — broken vs fixed OGG pair (same text/seed, only the reference encoder changed), corrected story 1, the three real reference recordings (3.1 s champú, 12.5 s, 30.4 s), and the 12-vs-30 generated A/B.
  • scripts/ — the exact harness used (story runners with cached/verified references, precision factorial, speaker scorer, reference builders, A/B tester).
Sample What to listen for
broken_awq_joinedref_int8enc.ogg vs fixed_awq_joinedref_fp32enc.ogg Same text/seed; only reference encoder differs
fixed_story1_fp32ref.ogg Corrected full story (FP32 ref, AWQ gen, INT8 decode)
reference_original_champu_003449.ogg / ref_2121_verified_12s_v2.ogg / ref_2121_verified_30s.ogg Real reference recordings, not generated speech
A_reference_12s.ogg / B_reference_30s.ogg Reference-length A/B, same text/seed/prompt

Reproduce

## corrected full stories (idle GPU 2, frozen env)
CUDA_VISIBLE_DEVICES=2 PYTORCH_ALLOC_CONF=expandable_segments:True \
  WORKSTATION/python run_stories_int8codec.py \
  --reference WORKSTATION/argentina-female-2121-003449.wav \
  --reference-encoder fp32 --prefix fp32ref_ \
  --outdir out_clone2121_fp32ref --metrics logs/stories_clone2121_fp32ref.json

## factorial isolation of generator/encoder/decoder
python compare_clone_precision.py          # -> precision_controls/results.json

## speaker-embedding evaluation (8 s windows, 4 s hop)
python score_clone_speakers.py --new-only --full-coverage --output speaker_scores_new.json

Models/paths are workstation-local: $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4, $WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5 (BF16 control), $WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1 (INT8 codec), speaker model microsoft/wavlm-base-plus-sv (HF cache; HF_HUB_DISABLE_XET=1).

Boundaries

Measured-result vs interpretation discipline applies (see repo README). WavLM cosine is a local speaker-resemblance proxy — not an accent, gender, or calibrated identity test; listening acceptance is tracked separately and the accent issue remains open. Single-pair A/Bs, single seed, Marlin/bf16 nondeterminism caveat. The INT8 encoder’s offending layers were not isolated; the fix bypasses that path.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS