Journal

AWQ speech: fast enough, with a memory catch

AWQ passed 33/33 prompts with pooled real-time factor 0.43 and 35.8 tokens per second.

Vlad / experimentos.
The boundary

The 13.05-GiB decode peak blocks a direct 12GB deployment. Style and listening questions remain separate from pass counts.

The question

AWQ speech delivered a useful render rate, but fitting a smaller GPU was still a separate question. The benchmark kept prompt health, pooled real-time factor, decode speed and peak memory in the same record.

From the original notebook

Benchmark + samples for the 4-bit AWQ build (lemuriandezapada/MOSS-TTS-v1.5-awq-int4, 6.3 GB) against the BF16 and HQ4 baselines. Same 33 prompts, seeds (1234) and sampling throughout; Marlin/bf16 kernels, SDPA, batch 1, one RTX 3090.

Code: groxaxo/moss-tts-hq4-quant/awq-3090 (full harness, env pins, reproduction commands).

Headline

  • 33/33 pass, 0 silent — pooled RTF 0.43 (BF16 0.59, HQ4 0.88), 35.8 tok/s (BF16 25.5, HQ4 17.0), peak VRAM 13.05 GiB combined.
  • Spanish: all 8 prompts healthy, no truncation/silence (see samples).
  • Register control (instruction-guided female, F0 ≥195 Hz): 3/15 (HQ4 1/21, BF16 1/4) — usable control restored, both languages pass.
  • Caveats: Marlin bf16 is run-to-run nondeterministic on sm86 (treat rates as single-draw estimates); a gptqmodel unload-retention quirk keeps decode peak at ~13 GiB (12GB deployment blocked until resolved).

Contents

  • docs/ — comparison.json (BF16/HQ4/AWQ), benchmark_awq.json (per-prompt records), register reports, single_run.json, FINDINGS.md
  • samples/ — smoke (Spanish), 3 Spanish suite prompts, 3 register passes, plus the 33-prompt prompts.txt
Sample What
awq3060-single.wav Smoke test, Spanish
es-conv-1.wav / es-long-1.wav / longform-es-1.wav Spanish short → 58 s longform
register-en-3.wav (203.4 Hz) / register-es-3.wav (198.3 Hz) Seeds-1235 female passes
register-es-3-seeds501.wav (219.2 Hz) Seeds-501 female pass (same seed HQ4 barely passed at 196.7)

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS