Back to the experimentSupporting notebook

NF4 deployment — polish-400 on RTX 3060 (mfwin) + Hermes default TTS

Source: breeze-tts2/latam-es-mf-polish-2026-09/docs/nf4-deploy.md · revision 6550ead3945b

Deployed the promoted polish-400 adapter as a merged model with bitsandbytes NF4 on the Windows GPU box (mfwin, research workstation) inside WSL2 Ubuntu 24.04, served over an OpenAI-compatible FastAPI /v1/audio/speech endpoint, and set as the default TTS (breeze_latam) for Hermes on every machine on the tailnet that runs Hermes (.51, .54, .55, mac).

Artifacts in this folder

Path What
docs/nf4-deploy.md This report
docs/nf4-bench-10.json Bench + WER JSON for both 10-prompt configs
nf4/tts10_*.wav 10 sentences from the winning config (B, temp 0.7)
nf4/prompts10.txt Their reference texts
deploy/ Repro kit: infer_nf4.py, server.py, breeze-tts.service, hermes_breeze_tts.py, fleet_breeze_cfg.py, wer_nf4.py

Model (7.2 GB, not in git): merged polish-400, built locally with merge_polish_400.py (adapter-vs-merged loss 6.6376 vs 6.6384 over 16 validation examples + fresh-reload check, receipt merge-400-receipt.json).

Remote layout (WSL op on mfwin)

WORKSTATION/breeze-tts2-nf4
  .venv/         torch 2.9.1+cu126, transformers 4.57.3, qwen-tts 0.1.1,
                 bitsandbytes 0.50.2, accelerate, fastapi/uvicorn
  merged-400/    merged model (sha256-verified after transfer)
  server.py      FastAPI app (systemd service `breeze-tts`)
  infer_nf4.py   CLI inference + benchmark (defaults: backbone-depth, temp 0.7)

Why WSL: bitsandbytes has no official Windows build. The WSL instance is itself the op-2 tailnet node (research workstation, tailscaled via systemd), so the server binds 0.0.0.0:12439 and is directly reachable fleet-wide — no Windows portproxy needed.

Quantization (tested default)

Benchmarks (RTX 3060, xAI STT judge es)

5 fixed prompts:

config WER RTF peak VRAM
release adapter-BF16 (3090, ref) 6.1% (3/49) — —
adapter-BF16, seeds 42–46 14.3% — —
merged-BF16, seeds 42–46 18.4% — —
merged-NF4, seeds 42–46, t0.9 24.5% ~6.9 5.49 GB
merged-NF4, seeds 100–104, t0.9 16.3% ~6.9 5.49 GB
merged-NF4, seeds 42–46, t0.7 16.3% ~6.9 5.49 GB

10 new prompts (nf4/prompts10.txt, temp 0.7, seeds 42–51):

config WER RTF peak VRAM
A: backbone-only NF4 (depth BF16) 29.9% (29/97) ~5.6 6.01 GB
B: backbone+depth NF4 (shipped) 24.7% (24/97) ~6.9 5.49 GB

Seed noise (±8 pts at temp 0.9) dominates the merge/NF4 deltas; A is faster but worse and sits 20 MB under the VRAM ceiling. TTFT ≈ 1.2 s, load ≈ 8 s.

Server API (OpenAI-compatible)

Autostart (boot chain, all verified except the reboot itself)

  1. Windows Task Scheduler BreezeTTS-BootWSL: ONSTART, SYSTEM, highest — wsl.exe -d Ubuntu /usr/bin/systemctl is-system-running --wait (created + ran on demand successfully).
  2. WSL systemd (enabled): tailscaled + breeze-tts.service.
  3. vmIdleTimeout=-1 keeps WSL alive once booted.

Hermes integration