Back to the experimentSupporting notebook

MOSS-TTS AWQ + INT8 codec: cloning failure isolated to the INT8 reference encoder (2026-09-21)

Source: moss-tts-awq/clone-fp32-reference-2026-09-21/README.md · revision 6550ead3945b

Zero-shot voice cloning with the AWQ INT4 generator and the HQ4-I8-SAFE codec produced the wrong voice (male) from a female Argentine reference. A controlled factorial comparison (generator × reference encoder × output decoder) isolated the cause: the INT8 codec was also encoding the reference audio, corrupting the conditioning tokens. Encoding the reference in FP32 and keeping INT8 only for output decoding restored target-speaker similarity from 0.8475 → 0.9556 (AWQ) and 0.8222 → 0.9591 (BF16 generator).

Working directory with live artifacts: $WORKSTATION/moss-awq4-stories/ (reports: CLONING-REVIEW.md, CLONING-FIX-FP32-REFERENCE.md, PROMPT-ACCENT-REVIEW.md, REFERENCE-12S-REVIEW.md, REFERENCE-30S-REVIEW.md, TEST-REPORT.md).

Headline

Contents

Sample What to listen for
broken_awq_joinedref_int8enc.ogg vs fixed_awq_joinedref_fp32enc.ogg Same text/seed; only reference encoder differs
fixed_story1_fp32ref.ogg Corrected full story (FP32 ref, AWQ gen, INT8 decode)
reference_original_champu_003449.ogg / ref_2121_verified_12s_v2.ogg / ref_2121_verified_30s.ogg Real reference recordings, not generated speech
A_reference_12s.ogg / B_reference_30s.ogg Reference-length A/B, same text/seed/prompt

Reproduce

## corrected full stories (idle GPU 2, frozen env)
CUDA_VISIBLE_DEVICES=2 PYTORCH_ALLOC_CONF=expandable_segments:True \
  WORKSTATION/python run_stories_int8codec.py \
  --reference WORKSTATION/argentina-female-2121-003449.wav \
  --reference-encoder fp32 --prefix fp32ref_ \
  --outdir out_clone2121_fp32ref --metrics logs/stories_clone2121_fp32ref.json

## factorial isolation of generator/encoder/decoder
python compare_clone_precision.py          # -> precision_controls/results.json

## speaker-embedding evaluation (8 s windows, 4 s hop)
python score_clone_speakers.py --new-only --full-coverage --output speaker_scores_new.json

Models/paths are workstation-local: $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4, $WORKSTATION/MOSS-TTS/weights/MOSS-TTS-v1.5 (BF16 control), $WORKSTATION/MOSS-TTS/experiments/hq4-int8/artifacts/codec_safe_v1 (INT8 codec), speaker model microsoft/wavlm-base-plus-sv (HF cache; HF_HUB_DISABLE_XET=1).

Boundaries

Measured-result vs interpretation discipline applies (see repo README). WavLM cosine is a local speaker-resemblance proxy — not an accent, gender, or calibrated identity test; listening acceptance is tracked separately and the accent issue remains open. Single-pair A/Bs, single seed, Marlin/bf16 nondeterminism caveat. The INT8 encoder’s offending layers were not isolated; the fix bypasses that path.