Back to the experimentSupporting notebook

Latam-ES Breeze TTS2 — M+F polish (2026-09)

Source: breeze-tts2/latam-es-mf-polish-2026-09/JOURNEY.md · revision 6550ead3945b

Journey (do not ship final step)

  1. Female LoRA — initial Latam-ES female-only adapter.
  2. M+F base run — continued to step 10250 (latam-es-mf).
  3. Paso0 WER screen — best base checkpoint by WER was checkpoint-step-008000 (not the final 10250).
  4. Polish — LoRA polish started from adapter at step 8000, run latam-es-mf-polish, 1500 polish steps on cuda:0, checkpoints every 100 steps.
  5. Polish WER (xAI STT, es) over the 5 fixed Spanish prompts:
    • Best: checkpoint-step-000400 @ 6.1% WER (3/49 edits)
    • Final checkpoint-step-001500 @ 42.9% WER — worse; do not ship.
  6. Release choice — promote polish-400 only; other polish checkpoints deleted on disk after selection.

Release artifacts here

Path What
samples/sample_01.wav … sample_05.wav Spanish samples from polish-400
samples/prompts.txt Matching reference prompts
docs/wer_report.md / wer_report.json Full WER table 100→1500
docs/polish-run-config.json Polish train config
docs/latest.json Points at promoted step 400

Polish params (summary)

See docs/polish-run-config.json for full config. High level:

WER ranking (avg over 5 prompts)

step avg WER
400 6.1%
100 10.2%
900 12.2%
200 / 1100 14.3%
1000 / 1200 16.3%
700 18.4%
800 24.5%
300 / 600 26.5%
1300 / 1400 28.6%
500 / 1500 42.9%

STT judge: local xAI audio shim → grok-transcribe, language es.

Explicit non-goals

Deployment (2026-09-11)

Polish-400 merged + NF4 on the mfwin RTX 3060 (WSL2), served via OpenAI-compatible POST /v1/audio/speech, set as Hermes default TTS (breeze_latam) on .51/.54/.55/mac with boot autostart. Full report: docs/nf4-deploy.md. Winner 10-prompt wavs under nf4/; reproducible deploy kit under deploy/.