Back to the experimentSupporting notebook

moss54d auto-chunk (2026-09-21)

Source: moss-tts-hq4/moss54d-autochunk-2026-09-21/README.md · revision 6550ead3945b

Sentence-aware auto-chunking for the live MOSS-TTS v1.5 HQ4-AWQ stack (daemon moss54d on :8894 + golden CLI infer_hybrid_3060.py). Long inputs split on sentence boundaries (dots, ?, !, …, newlines; abbreviations like Sr./EE.UU. protected); each chunk renders separately with bounded VRAM and joins with 150 ms of silence.

Snapshot of the exact files deployed to WORKSTATION/moss54-staging on host research workstation (RTX 3060 12 GB). Source of truth stays in staging; this dir is the reproducible-experiment copy (NOT mirrored to vladgateway).

Files

Semantics

Verification (live daemon, 2026-09-21)

Case chunks Audio Tokens Total Health
short (Hola, esta es una prueba corta…) 1 4.8 s 94 12.9 s rms 0.095, peak 0.86, not silent
long (551 chars, 8 sentences) 2 ([282, 269] chars) 48.0 s 666 92.5 s rms 0.078, peak 0.84, not silent

Daemon log: [moss54d] auto-chunk: 2 chunks ([282, 269] chars). GPU steady at ~9.6 GiB (vLLM 2.5 + daemon 6.4 + headroom) — no VRAM explosion. samples/ holds both wavs rendered with seed=1234, language=Spanish.

Existing suite tests/test_runtime_guards.py: 5/5 pass after the change.

Note: the golden CLI chunk loop is code-complete and shares the unit-tested splitter, but a live CLI long-render was not run (it would exceed 12 GB alongside the warm daemon — needs a maintenance window).