Journal

Long speech, bounded memory

A 551-character, eight-sentence input rendered in two chunks with 48.0 seconds of healthy audio.

Vlad / experimentos.
The boundary

The live daemon was verified. A separate live long-render through the CLI was deferred because of memory contention.

The question

Long input needed a bounded rendering path. The live service split an eight-sentence request into chunks and checked the resulting audio, while documenting the separate CLI run that remained deferred.

From the original notebook

Sentence-aware auto-chunking for the live MOSS-TTS v1.5 HQ4-AWQ stack (daemon moss54d on :8894 + golden CLI infer_hybrid_3060.py). Long inputs split on sentence boundaries (dots, ?, !, …, newlines; abbreviations like Sr./EE.UU. protected); each chunk renders separately with bounded VRAM and joins with 150 ms of silence.

Snapshot of the exact files deployed to WORKSTATION/moss54-staging on host research workstation (RTX 3060 12 GB). Source of truth stays in staging; this dir is the reproducible-experiment copy (NOT mirrored to vladgateway).

Files

  • chunking.py — pure-python splitter: split_sentences, chunk_text (default 300 chars/chunk, hard-split fallback for dotless text).
  • moss54d.py — warm daemon; speech() dispatches to _render_single per chunk. New request fields: auto_chunk (default true), chunk_chars, chunk_gap_ms. Response adds "chunks".
  • infer_hybrid_3060.py — golden CLI with the same chunk loop inside run_single (model stays loaded across chunks). Flags: --no-auto-chunk, --chunk-chars, --chunk-gap-ms.

Semantics

  • 1 chunk → byte-identical to the old single-pass path (same seed).
  • N chunks → chunk i renders with seed+i (deterministic on both paths).
  • Per-chunk audio-health gate (error names the failing chunk).
  • Reference voice (--reference / voice) reused for every chunk.

Verification (live daemon, 2026-09-21)

Case chunks Audio Tokens Total Health
short (Hola, esta es una prueba corta…) 1 4.8 s 94 12.9 s rms 0.095, peak 0.86, not silent
long (551 chars, 8 sentences) 2 ([282, 269] chars) 48.0 s 666 92.5 s rms 0.078, peak 0.84, not silent

Daemon log: [moss54d] auto-chunk: 2 chunks ([282, 269] chars). GPU steady at ~9.6 GiB (vLLM 2.5 + daemon 6.4 + headroom) — no VRAM explosion. samples/ holds both wavs rendered with seed=1234, language=Spanish.

Existing suite tests/test_runtime_guards.py: 5/5 pass after the change.

Note: the golden CLI chunk loop is code-complete and shares the unit-tested splitter, but a live CLI long-render was not run (it would exceed 12 GB alongside the warm daemon — needs a maintenance window).

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS