Back to the experimentSupporting notebook

NF4 re-host on op (.54) — lazy loading + fleet re-point (2026-09-13)

Source: breeze-tts2/latam-es-mf-polish-2026-09/docs/nf4-rehost-op54-lazy.md · revision 6550ead3945b

Moved the default TTS endpoint off mfwin-WSL (op-2, research workstation) to op (4790, research workstation):12439 with lazy loading, per request. The op-2 service keeps running as fallback (mac still points at it until it wakes).

Layout on .54 (WORKSTATION/breeze-tts2-nf4)

.venv/         torch 2.9.1+cu126, transformers 4.57.3, qwen-tts 0.1.1,
               bitsandbytes 0.50.2, accelerate 1.12.0, fastapi/uvicorn
merged-400/    staged from .51 (sha256-verified, see receipt below)
breeze_infer/ models/ training/   code rsynced from breeze-tts2-finetuning
server.py      deploy/server_lazy.py (lazy variant, see below)
breeze-tts.service  systemd unit (same as deploy/, After=tailscaled)

Source of the model: .51:WORKSTATION/merged-400 (op-2 has no sshd by design; mfwin:22 closed — .51 had the merge output).

Lazy loading (deploy/server_lazy.py)

Same generation path as deploy/server.py (NF4 backbone+depth, temp 0.7, seed 42, silence early-stop + trim, wav/mp3/pcm). Differences:

Rationale: .54 shares the RTX 3060 with the MemPalace vLLM (Nemotron AWQ, ~2.5 GB) — don’t grab 5.5 GB at boot or on every restart.

Verification receipts

Rollback

Per machine: restore config.yaml.bak-breeze54-20260913. On .54: sudo systemctl disable --now breeze-tts (frees 5.4 GB VRAM; vLLM was never touched). op-2 endpoint stays up regardless.