Journal

The best checkpoint was not the last one

Polish step 400 measured 6.1% word error rate; the final step 1500 measured 42.9%.

Vlad / experimentos.
The boundary

The screen uses five fixed Spanish prompts and an STT judge. Word error rate does not measure accent, timbre or listener preference.

Measurements

Recorded result

The best checkpoint was not the last one

Five fixed Spanish prompts · 49 reference words · STT screen

The best checkpoint was not the last oneStep 100: 10.2 word error rate (%); Step 200: 14.3 word error rate (%); Step 300: 26.5 word error rate (%); Step 400: 6.1 word error rate (%); Step 500: 42.9 word error rate (%); Step 600: 26.5 word error rate (%); Step 700: 18.4 word error rate (%); Step 800: 24.5 word error rate (%); Step 900: 12.2 word error rate (%); Step 1000: 16.3 word error rate (%); Step 1100: 14.3 word error rate (%); Step 1200: 16.3 word error rate (%); Step 1300: 28.6 word error rate (%); Step 1400: 28.6 word error rate (%); Step 1500: 42.9 word error rate (%). Transcript error does not measure accent, timbre or listener preference.0%12%24%36%48%Step 100: 10.2%Step 200: 14.3%Step 300: 26.5%Step 400: 6.1%Best: 6.1%Step 500: 42.9%Step 600: 26.5%Step 700: 18.4%Step 800: 24.5%Step 900: 12.2%Step 1000: 16.3%Step 1100: 14.3%Step 1200: 16.3%Step 1300: 28.6%Step 1400: 28.6%Step 1500: 42.9%10050010001500Training step · word error rate (%)
0%24%48%Step 100: 10.2%Step 200: 14.3%Step 300: 26.5%Step 400: 6.1%Step 500: 42.9%Step 600: 26.5%Step 700: 18.4%Step 800: 24.5%Step 900: 12.2%Step 1000: 16.3%Step 1100: 14.3%Step 1200: 16.3%Step 1300: 28.6%Step 1400: 28.6%Step 1500: 42.9%10050010001500Training step

Transcript error does not measure accent, timbre or listener preference.

View data & source
The best checkpoint was not the last one · word error rate (%)
ConfigurationValue
Step 10010.2
Step 20014.3
Step 30026.5
Step 4006.1
Step 50042.9
Step 60026.5
Step 70018.4
Step 80024.5
Step 90012.2
Step 100016.3
Step 110014.3
Step 120016.3
Step 130028.6
Step 140028.6
Step 150042.9

Origin: structured measurements. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

Training longer did not make every speech checkpoint better. The fixed Spanish transcription screen compared the saved steps, found an earlier best checkpoint, and retained the distinction between word accuracy and how a voice sounds.

From the original notebook

Journey (do not ship final step)

  1. Female LoRA — initial Latam-ES female-only adapter.
  2. M+F base run — continued to step 10250 (latam-es-mf).
  3. Paso0 WER screen — best base checkpoint by WER was checkpoint-step-008000 (not the final 10250).
  4. Polish — LoRA polish started from adapter at step 8000, run latam-es-mf-polish, 1500 polish steps on cuda:0, checkpoints every 100 steps.
  5. Polish WER (xAI STT, es) over the 5 fixed Spanish prompts:
    • Best: checkpoint-step-000400 @ 6.1% WER (3/49 edits)
    • Final checkpoint-step-001500 @ 42.9% WER — worse; do not ship.
  6. Release choice — promote polish-400 only; other polish checkpoints deleted on disk after selection.

Release artifacts here

Path What
samples/sample_01.wav … sample_05.wav Spanish samples from polish-400
samples/prompts.txt Matching reference prompts
docs/wer_report.md / wer_report.json Full WER table 100→1500
docs/polish-run-config.json Polish train config
docs/latest.json Points at promoted step 400

Polish params (summary)

See docs/polish-run-config.json for full config. High level:

  • Base adapter init: WORKSTATION/checkpoint-step-008000
  • Polish run root: WORKSTATION/latam-es-mf-polish
  • Device: cuda:0 (power-capped card; intentional for this polish)
  • Steps: 1500; save every 100
  • Promoted checkpoint on server: WORKSTATION/checkpoint-step-000400

WER ranking (avg over 5 prompts)

step avg WER
400 6.1%
100 10.2%
900 12.2%
200 / 1100 14.3%
1000 / 1200 16.3%
700 18.4%
800 24.5%
300 / 600 26.5%
1300 / 1400 28.6%
500 / 1500 42.9%

STT judge: local xAI audio shim → grok-transcribe, language es.

Explicit non-goals

  • Do not ship polish-1500 / TRAIN_DONE final solely because training completed.
  • Prefer polish-400 (or re-listen top-3: 400, 100, 900) for release.

Deployment (2026-09-11)

Polish-400 merged + NF4 on the mfwin RTX 3060 (WSL2), served via OpenAI-compatible POST /v1/audio/speech, set as Hermes default TTS (breeze_latam) on .51/.54/.55/mac with boot autostart. Full report: docs/nf4-deploy.md. Winner 10-prompt wavs under nf4/; reproducible deploy kit under deploy/.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (5)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS