Deployed the promoted polish-400 adapter as a merged model with bitsandbytes
NF4 on the Windows GPU box (mfwin, research workstation) inside WSL2 Ubuntu 24.04,
served over an OpenAI-compatible FastAPI /v1/audio/speech endpoint, and set
as the default TTS (breeze_latam) for Hermes on every machine on the
tailnet that runs Hermes (.51, .54, .55, mac).
Artifacts in this folder
| Path | What |
|---|---|
docs/nf4-deploy.md |
This report |
docs/nf4-bench-10.json |
Bench + WER JSON for both 10-prompt configs |
nf4/tts10_*.wav |
10 sentences from the winning config (B, temp 0.7) |
nf4/prompts10.txt |
Their reference texts |
deploy/ |
Repro kit: infer_nf4.py, server.py, breeze-tts.service, hermes_breeze_tts.py, fleet_breeze_cfg.py, wer_nf4.py |
Model (7.2 GB, not in git): merged polish-400, built locally with
merge_polish_400.py (adapter-vs-merged loss 6.6376 vs 6.6384 over 16
validation examples + fresh-reload check, receipt merge-400-receipt.json).
Remote layout (WSL op on mfwin)
WORKSTATION/breeze-tts2-nf4
.venv/ torch 2.9.1+cu126, transformers 4.57.3, qwen-tts 0.1.1,
bitsandbytes 0.50.2, accelerate, fastapi/uvicorn
merged-400/ merged model (sha256-verified after transfer)
server.py FastAPI app (systemd service `breeze-tts`)
infer_nf4.py CLI inference + benchmark (defaults: backbone-depth, temp 0.7)
Why WSL: bitsandbytes has no official Windows build. The WSL instance is
itself the op-2 tailnet node (research workstation, tailscaled via systemd), so the
server binds 0.0.0.0:12439 and is directly reachable fleet-wide — no
Windows portproxy needed.
Quantization (tested default)
- NF4 double-quant (
bnb_4bit_compute_dtype=bf16) onbackbone_model+depth_decoder(282 linears);codec_model,text_encoder, embeddings,lm_headstay BF16. Temperature default 0.7, seed 42 (voice: latam). - Two gotchas (both bitten, both fixed in the shipped code):
- transformers 4.57 has no
modules_to_not_convertonBitsAndBytesConfig; the 4-bit quantizer reads its skip list fromllm_int8_skip_modules. Wrong kwarg = silently quantizes everything. from_pretrainedmust passdtype=torch.bfloat16, else kept modules load in FP16 and the T5 text encoder overflows (inf → NaN → CUDA assert in sampling).
- transformers 4.57 has no
- Server additions beyond the CLI: trailing-silence early-stop (2 s) + leading/trailing trim (150 ms pads). The model often fails to emit EOS and pads with silence (e.g. 25 s out for a 3 s sentence); trimmed output transcribes identically (verified by STT).
Benchmarks (RTX 3060, xAI STT judge es)
5 fixed prompts:
| config | WER | RTF | peak VRAM |
|---|---|---|---|
| release adapter-BF16 (3090, ref) | 6.1% (3/49) | — | — |
| adapter-BF16, seeds 42–46 | 14.3% | — | — |
| merged-BF16, seeds 42–46 | 18.4% | — | — |
| merged-NF4, seeds 42–46, t0.9 | 24.5% | ~6.9 | 5.49 GB |
| merged-NF4, seeds 100–104, t0.9 | 16.3% | ~6.9 | 5.49 GB |
| merged-NF4, seeds 42–46, t0.7 | 16.3% | ~6.9 | 5.49 GB |
10 new prompts (nf4/prompts10.txt, temp 0.7, seeds 42–51):
| config | WER | RTF | peak VRAM |
|---|---|---|---|
| A: backbone-only NF4 (depth BF16) | 29.9% (29/97) | ~5.6 | 6.01 GB |
| B: backbone+depth NF4 (shipped) | 24.7% (24/97) | ~6.9 | 5.49 GB |
Seed noise (±8 pts at temp 0.9) dominates the merge/NF4 deltas; A is faster but worse and sits 20 MB under the VRAM ceiling. TTFT ≈ 1.2 s, load ≈ 8 s.
Server API (OpenAI-compatible)
POST /v1/audio/speech{model?, input, voice?, response_format? (wav/mp3/pcm), speed? (=1.0 only)}plus extrastemperature?, seed?, instruction?, cfg_scale?, repetition_penalty?.input1–600 chars. One generation at a time (GPU lock); extras ignored by strict OpenAI clients.GET /v1/models,GET /health(ready/vram/queue),GET /.- Verified: wav + mp3 output, empty/too-long/bad-format/bad-speed rejections,
trimmed-content STT parity, follow-up request health after early-stop,
crash auto-restart (systemd
Restart=always, back in <30 s).
Autostart (boot chain, all verified except the reboot itself)
- Windows Task Scheduler
BreezeTTS-BootWSL: ONSTART, SYSTEM, highest —wsl.exe -d Ubuntu /usr/bin/systemctl is-system-running --wait(created + ran on demand successfully). - WSL systemd (enabled):
tailscaled+breeze-tts.service. vmIdleTimeout=-1keeps WSL alive once booted.
Hermes integration
- New command provider
breeze_latam(hermes_breeze_tts.py→ the endpoint),tts.provider: breeze_latamon .51, .54, .55, mac (backupsconfig.yaml.bak-breeze-20260911next to each config; no gateway restart needed — Hermes invalidates config cache on mtime change). - Verified per machine: direct wrapper synthesis on all four; full
text_to_speech_toolpath on .51 (provider: breeze_latam, ogg voice message out) and .54. Previous providers kept as fallback (chatterbox_iris / indextts-2-5 / supertonic / neutts_air_q8). - Caps: provider
max_text_length: 500, wrapper/server timeouts 600/660 s. Previous default on .51 waschatterbox_iris.