Host: research workstation, GPU 2 (RTX 3090 24GB).
Replaces the MOSS-TTS v1.5 HQ4 + FP32-codec server with HQ4 + full INT8-RTN
codec behind a lazy gate. Public contract unchanged (moss-8b,
POST /v1/audio/speech → WAV on :9840).
What
- INT8-RTN codec master (
weight_int8I8 + per-channel F32 scale, 620 modules) staged at$WORKSTATION/MOSS-TTS/weights/MOSS-Audio-Tokenizer-INT8. - Server dequantizes inline at load: 554 linears → BF16, 66 weight-norm convs → exact F32 g/v re-split, codebooks stay F32. Strict key check (1600/1600 vs FP32 reference); no-op for FP32/BF16 dirs.
- Codec encode/decode run under BF16 autocast when INT8 is active (fixes
the
float vs BFloat16matmul mismatch; FP32 path bit-identical). - Backend moved to loopback
localhost, autostart off; lazy gate on:9840starts it on first request, reaps after 300s idle. - HQ4 weights reused in place: safetensors sha256-verified identical to the published artifact (Hub repo is private, no token available).
Measured
- Resident: ~10.3GB (was ~13.9GB FP32).
- INT8 renders healthy (male/short/female-clone); gate cold start 14.8s first render, 5.7s warm; reap + re-wake verified.
- Note: 7-char prompts sometimes render mostly silence on FP32 and INT8 alike (sampling variance, pre-existing).
Files
moss_server_cuda.py— server (deployed to$WORKSTATION/moss-server/).moss-tts-cuda.service— backend unit (loopback:17790).moss-tts-lazy-proxy.service— gate unit (:9840→:17790).dequant_codec_int8_to_bf16.py— provenance tool: builds a loadable BF16 runtime from the INT8 master (unused at runtime; server dequants inline because this host’s disk was full).verify_int8_rt.py— CPU check of the server’s exact codec load path.
Redeploy
cp moss_server_cuda.py $WORKSTATION/moss-server/moss_server_cuda.py
cp moss-tts-cuda.service moss-tts-lazy-proxy.service $WORKSTATION/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user disable moss-tts-cuda.service
systemctl --user enable --now moss-tts-lazy-proxy.service