Journal

A speech service that sleeps between requests

The recorded service used about 10.3GB resident memory; first render took 14.8 seconds and warm render 5.7.

Vlad / experimentos.
The boundary

Cold and warm timings describe different service states. Short-prompt silence existed in both FP32 and INT8 paths.

Measurements

Recorded result

First request and warm request are different states

HQ4 + INT8 codec lazy service · recorded request timings

First request and warm request are different statesCold first render: 14.8 seconds; Warm render: 5.7 seconds. Cold start includes service wake-up. This is not two equivalent decoding workloads.Cold first render14.8Cold first render: 14.8 secondsWarm render5.7Warm render: 5.7 seconds0seconds
  1. Cold first render14.8
  2. Warm render5.7

seconds

Cold start includes service wake-up. This is not two equivalent decoding workloads.

View data & source
First request and warm request are different states · seconds
ConfigurationValue
Cold first render14.8
Warm render5.7

Origin: reported live-service timings. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

A speech service need not hold all of its resources between requests. The lazy-server experiment checked wake-up, warm rendering, reap and re-wake, while preserving the known short-prompt silence behavior.

From the original notebook

Host: research workstation, GPU 2 (RTX 3090 24GB).

Replaces the MOSS-TTS v1.5 HQ4 + FP32-codec server with HQ4 + full INT8-RTN codec behind a lazy gate. Public contract unchanged (moss-8b, POST /v1/audio/speech → WAV on :9840).

What

  • INT8-RTN codec master (weight_int8 I8 + per-channel F32 scale, 620 modules) staged at $WORKSTATION/MOSS-TTS/weights/MOSS-Audio-Tokenizer-INT8.
  • Server dequantizes inline at load: 554 linears → BF16, 66 weight-norm convs → exact F32 g/v re-split, codebooks stay F32. Strict key check (1600/1600 vs FP32 reference); no-op for FP32/BF16 dirs.
  • Codec encode/decode run under BF16 autocast when INT8 is active (fixes the float vs BFloat16 matmul mismatch; FP32 path bit-identical).
  • Backend moved to loopback localhost, autostart off; lazy gate on :9840 starts it on first request, reaps after 300s idle.
  • HQ4 weights reused in place: safetensors sha256-verified identical to the published artifact (Hub repo is private, no token available).

Measured

  • Resident: ~10.3GB (was ~13.9GB FP32).
  • INT8 renders healthy (male/short/female-clone); gate cold start 14.8s first render, 5.7s warm; reap + re-wake verified.
  • Note: 7-char prompts sometimes render mostly silence on FP32 and INT8 alike (sampling variance, pre-existing).

Files

  • moss_server_cuda.py — server (deployed to $WORKSTATION/moss-server/).
  • moss-tts-cuda.service — backend unit (loopback :17790).
  • moss-tts-lazy-proxy.service — gate unit (:9840 → :17790).
  • dequant_codec_int8_to_bf16.py — provenance tool: builds a loadable BF16 runtime from the INT8 master (unused at runtime; server dequants inline because this host’s disk was full).
  • verify_int8_rt.py — CPU check of the server’s exact codec load path.

Redeploy

cp moss_server_cuda.py $WORKSTATION/moss-server/moss_server_cuda.py
cp moss-tts-cuda.service moss-tts-lazy-proxy.service $WORKSTATION/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user disable moss-tts-cuda.service
systemctl --user enable --now moss-tts-lazy-proxy.service

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS