Back to the experimentSupporting notebook

INT8 codec + lazy CUDA server (2026-09-19)

Source: moss-tts-hq4/int8-codec-lazy-server-2026-09-19/README.md · revision 6550ead3945b

Host: research workstation, GPU 2 (RTX 3090 24GB).

Replaces the MOSS-TTS v1.5 HQ4 + FP32-codec server with HQ4 + full INT8-RTN codec behind a lazy gate. Public contract unchanged (moss-8b, POST /v1/audio/speech → WAV on :9840).

What

Measured

Files

Redeploy

cp moss_server_cuda.py $WORKSTATION/moss-server/moss_server_cuda.py
cp moss-tts-cuda.service moss-tts-lazy-proxy.service $WORKSTATION/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user disable moss-tts-cuda.service
systemctl --user enable --now moss-tts-lazy-proxy.service