Moved the default TTS endpoint off mfwin-WSL (op-2, research workstation) to op (4790, research workstation):12439 with lazy loading, per request. The op-2 service keeps running as fallback (mac still points at it until it wakes).
Layout on .54 (WORKSTATION/breeze-tts2-nf4)
.venv/ torch 2.9.1+cu126, transformers 4.57.3, qwen-tts 0.1.1,
bitsandbytes 0.50.2, accelerate 1.12.0, fastapi/uvicorn
merged-400/ staged from .51 (sha256-verified, see receipt below)
breeze_infer/ models/ training/ code rsynced from breeze-tts2-finetuning
server.py deploy/server_lazy.py (lazy variant, see below)
breeze-tts.service systemd unit (same as deploy/, After=tailscaled)
Source of the model: .51:WORKSTATION/merged-400
(op-2 has no sshd by design; mfwin:22 closed — .51 had the merge output).
Lazy loading (deploy/server_lazy.py)
Same generation path as deploy/server.py (NF4 backbone+depth, temp 0.7,
seed 42, silence early-stop + trim, wav/mp3/pcm). Differences:
lifespandoes NOT load the model; service starts instantly with 0 VRAM.- First
POST /v1/audio/speechloads under the request lock (double-checked; concurrent callers wait), then it stays resident. GET /healthreportsstatus: idle | loading | readypluslazy,loads.- Knobs (env):
BREEZE_LAZY=0restores eager startup;BREEZE_MODEL_DIR;BREEZE_IDLE_UNLOAD_SEC>0optional idle unload (default off).
Rationale: .54 shares the RTX 3060 with the MemPalace vLLM (Nemotron AWQ, ~2.5 GB) — don’t grab 5.5 GB at boot or on every restart.
Verification receipts
- Model transfer:
sha256sum model-*.safetensorsidentical .51 vs .54 —457773cd…ac2ed64,1aff4236…8d76590. (First attempt nested the dir one level; flattened — tokenizer_class lookup failure was the symptom.) /health:{"status":"idle","lazy":true,"loads":0}at boot →readywithvram_allocated_gb: 5.442,nf4_modules: 282after first request. vLLM EngineCore (2 548 MiB) untouched throughout.- Speech: wav 4.24 s audio (seed 42) and mp3 128 kbps both 200 OK; first request ≈ 46 s incl. load, warm RTF ≈ 7.8 (GPU shared with vLLM; mfwin dedicated was ≈ 6.9).
- Fleet Hermes
breeze_latamURL →private workstation serviceon .51, .54, .55 (sed+ backupconfig.yaml.bak-breeze54-20260913; Hermes reloads on config mtime). Wrapper synthesis verified on each (deterministic wav, 105 644 B for the shared test line). - Pending: mac (op-1) asleep/unreachable — still on op-2:12439 (working fallback); flip when awake (same one-line sed).
Rollback
Per machine: restore config.yaml.bak-breeze54-20260913. On .54:
sudo systemctl disable --now breeze-tts (frees 5.4 GB VRAM; vLLM was never
touched). op-2 endpoint stays up regardless.