Back to the experimentSupporting notebook

MOSS-TTS-v1.5 AWQ INT4 on a single RTX 3090 — Findings

Source: moss-tts-awq/awq-3090-2026-09/docs/FINDINGS.md · revision 6550ead3945b

Scope note: the original task targeted an RTX 3060 12GB, but the lab machine has 3× RTX 3090 and no 3060. Per user direction the work was retargeted to one free RTX 3090 (physical GPU 2). File names keep the 3060 identifier for checklist compatibility; every measurement below is from the 3090 unless stated otherwise.

Hardware

Software versions

Package Version Notes
Python 3.11.16 (conda env moss-awq3060) conda create -n moss-awq3060 python=3.11 + pip stack
torch / torchaudio 2.9.1+cu128 PyTorch cu128 index, same as BF16 baseline
transformers 5.17.0 Baseline pins 5.0.0; 5.17 required by gptqmodel 7.5 (>=5.14)
gptqmodel 7.5.0 AWQ kernel backend (Marlin)
torchao 0.16.0+cu128 0.18 breaks on torch 2.9 (ScalingType import)
accelerate 1.15.0 Same as baseline
safetensors / huggingface_hub 0.8.0 / 1.32.0 Same as baseline
numpy 2.2.6 Baseline 2.1.0; 2.2.6 required by gptqmodel 7.5
torchcodec 0.8.1 (PyPI) +cu128 build fails to load libtorchcodec here (no system FFmpeg dev libs); PyPI build works
MOSS-TTS source OpenMOSS/MOSS-TTS @ 934d682, editable install --no-deps Cloned into MOSS-TTS/ inside the lab

Full freeze: logs/conda-pip-freeze.txt, logs/conda-list.txt. pip check deltas vs baseline: only the known moss-tts notes (gradio not installed, numpy/safetensors pins newer than pyproject).

Checkpoint

How AWQ was loaded

transformers-native AWQ path (AwqQuantizer, trust_remote_code=True, attn_implementation="sdpa", device_map="cuda:0"), kernels from gptqmodel 7.5.0, which auto-selects AwqMarlinLinear (Ampere Marlin; JIT-compiled once to $WORKSTATION/.cache/gptqmodel, ~104 s first run, cached after).

Two load-time interventions were required (see Failures/Fixes):

  1. bf16 compute override. transformers 5.17 force-casts AWQ to fp16, but Qwen3 outlier activations reach ~85504 (measured, decoder layer 7) and overflow fp16 (max 65504) → inf/NaN logits → CUDA device-side assert in sampling. gptqmodel Marlin supports bf16, so run_awq_3060.py monkeypatches AwqQuantizer.update_dtype to pass bf16 through. Model runs int4 weights / bf16 compute / bf16 untouched modules.
  2. No AutoAWQ package. transformers ≥5 loads AWQ via gptqmodel, not AutoAWQ; installing AutoAWQ 0.2.9 would have downgraded transformers to 4.x and broken the v1.5 custom modeling code. Verified: 252 quantized linears (AwqMarlinLinear), 33 plain Linear heads, 33 fp→bf16 embeddings.

Staged-loading architecture

src/run_awq_3060.py::run_single implements generate → full unload → decode:

  1. Processor on CPU; AWQ transformer on cuda:0 (codec parked on CPU during generation; reference-encode stages the codec briefly for clone prompts).
  2. Timed generation under torch.inference_mode() (no autocast — native bf16), peak stats reset around it.
  3. Token IDs → CPU; del model + CUDA tensors; gc.collect(); empty_cache(); ipc_collect(); synchronize; residual recorded.
  4. Codec → cuda:0 only afterwards; peak stats reset; timed decode → WAV.

Known wart: step 3 retains exactly one model’s VRAM (~6.24 GiB) due to an unidentified gptqmodel-side reference (see Failures). Decode peak therefore reads ~13.04 GiB (retained transformer + codec) instead of ~6.8 GiB codec-only. No unbounded accumulation (reload settles back to 6.24 GiB), and the 33-prompt benchmark uses a load-once design so this never compounds.

Single-prompt smoke result

VRAM measurements

(Single-prompt staged run, Marlin/bf16.)

Phase Peak allocated Peak reserved
Generation (transformer + KV) ~6.31 GiB ~6.38 GiB
Post-unload residual ~6.24 GiB ~6.32 GiB
Decode (retained transformer + codec) ~13.04 GiB ~13.10 GiB

Codec-only estimate: 13.04 − 6.24 ≈ 6.8 GiB. For comparison, BF16 peaks at 22.6 GiB combined (15.9 gen + codec) and HQ4 at ~13.2 GiB combined.

33-prompt benchmark

Result: 33/33 ok, 0 failed, 0 silent. Full per-prompt records: benchmarks/awq/benchmark_awq.json (+ results/benchmark_awq.json copy); WAVs in benchmarks/awq/wav/.

Metric (gen-only RTF, suite def.) BF16 HQ4 AWQ (Marlin/bf16)
Median RTF 0.59 0.96 0.47
Pooled RTF (Σgen/Σaudio) 0.59 0.88 0.43
Tokens/s (pooled) 25.5 17.0 35.8
Audio seconds 446.6 446.4 410.6
Compute seconds (gen) 262.5 394.5 174.5 (+3.8 decode)
Peak VRAM (suite max) 22.77 GiB 13.19 GiB 6.49 gen / 13.05 decode-combined
Weights on disk 16 GB 6.3 GB 6.3 GB

Honest end-to-end numbers (AWQ only; historical harness never timed decode): median e2e RTF 0.48, pooled e2e RTF 0.43 (decode adds ~2% — 3.8 s over 410.6 s audio). Gen-time p50 4.07 s, p95 12.15 s; model load 4.0 s (once, excluded from RTF). Median audio 7.76 s.

Spanish quality observations

All 8 Spanish-bearing prompts produced valid non-silent audio (es-conv-1, es-formal-1, es-question-1, es-numbers-1, es-long-1, mix-en-es-1, clone-es-1, longform-es-1): RMS 0.034–0.173, no NaN/Inf, sample rate 24 kHz throughout, ffprobe-clean. Durations track the baselines within normal sampling spread (e.g. longform-es-1: AWQ 58.0 s vs BF16 85.0 s vs HQ4 57.9 s; es-long-1: 17.4 s vs 11.7 / 26.9 s) — no systematic truncation (longform used 759 of 3000 budgeted tokens) and no silence or token-leakage signature in any Spanish waveform. No transcript/WER evaluator exists in the repos, so no WER claim is made; naturalness assessment beyond these signal checks needs blind listening on the paired WAVs.

Voice/register-control findings

Suite EN passes ES passes Combined
BF16 (seeds 1235) 1/1 (226.4 Hz) 0/3 (138–142 Hz) 1/4
HQ4 (1235 + 501) 0/12 1/9 (196.7 @503) 1/21
AWQ (1235 + 501) 1/9 (203.4 @1237) 2/6 (198.3 @1237, 219.2 @503) 3/15

AWQ restores usable instruction/register control where HQ4 had lost it (3× the passes in fewer attempts; female register reached in both languages), with a striking same-seed data point: seed 503/ES passes under both HQ4 (196.7 Hz, marginal) and AWQ (219.2 Hz, decisive). It does not fully match BF16’s first-try EN strength (226.4 Hz). Caveat: Marlin nondeterminism means a rerun draws different F0s; treat rates as single-draw estimates.

BF16 comparison

HQ4 comparison

Failures encountered

  1. No RTX 3060 on the machine (3× RTX 3090). Retargeted to one free 3090 per user direction; 3060-fit claims are projections, not measurements.
  2. fp16 activation overflow. transformers forces AWQ→fp16; Qwen3 outliers hit ~85504 at decoder layer 7 → inf/NaN logits → CUDA device-side assert in torch.multinomial at generate step 0. Localized with a forward-pass NaN probe (inf-first at layer 7, clean embeds).
  3. Marlin bf16 nondeterminism on sm86. Two identical in-process forwards differ by up to 23.0 in logits; same-seed generations diverge run to run (e.g. 174 vs 158 tokens). gptqmodel warns bf16-Marlin is a pre-SM90 compatibility path.
  4. Deterministic-backend bake-off failed on speed/correctness. TORCH_AWQ: deterministic, 1.9 tok/s (RTF 8.8). GEMM/bf16: deterministic, 3.1 tok/s (RTF 5.2). EXLLAMA_V2: deterministic but NaN logits. None viable as primary; Marlin kept as primary with documented caveat.
  5. Staged-unload retention. Post-generation del + gc + empty_cache retains exactly one model’s VRAM (6.24 GiB), backend-independent; decode peak reads ~13.04 GiB instead of ~6.8 GiB. No unbounded growth (reloads settle back to 6.24 GiB). Root holder unidentified after gc census, weakref census, referrer BFS, thread-stack walk, CUDA allocation snapshot, and inspection of custom modeling code, accelerate state, and gptqmodel post_init (no global caches found). Suspected invisible-to-gc reference in the quantizer load path. This is a 3060 blocker: per-item staged inference would peak ~19.5 GiB (stale + fresh + codec), exceeding 12 GB.
  6. Dependency churn. pip install gptqmodel silently upgraded transformers 5.0→5.17 and numpy 2.1→2.2.6 (detected via before/after freeze; kept deliberately — gptqmodel 7.5 requires transformers ≥5.14 and older gptqmodel has no usable wheel here). torchao 0.18 broke on torch 2.9 (ScalingType); pinned 0.16.0. torchcodec +cu128 couldn’t load libtorchcodec (no system FFmpeg libs); PyPI 0.8.1 works.
  7. venv → conda migration mid-task per user request; stack replicated pin-for-pin (106 packages) into moss-awq3060, smoke-verified.

Fixes applied

AWQ INT4 + Marlin/bf16 + SDPA, load-once, batch 1, on a 24GB card — faster than BF16, far less VRAM, better register control than HQ4:

conda run -n moss-awq3060 python src/benchmark_awq_3060.py  # or run_single.sh

Caveats: (a) nondeterministic outputs run-to-run — fine for serving, noisy for eval; (b) keep transformers==5.17.0 + gptqmodel==7.5.0 pinned as in logs/conda-pip-freeze.txt; (c) do NOT attempt 12GB deployment until the unload-retention issue is resolved (needs ≤6.8 GiB codec-only decode; currently ~13 GiB). Next optimization to try: Marlin fp16 with outlier taming (unlikely — outliers are structural), or exllamav3/qqq AWQ backends in newer gptqmodel for a deterministic fast path.

Exact reproduction commands

## 0. Prereqs: RTX 3090 GPU 2 free; miniforge3 conda; ffmpeg installed.
## 1. Env (already created; to recreate):
conda create -y -n moss-awq3060 python=3.11 pip
conda run -n moss-awq3060 pip install \
  --index-url https://download.pytorch.org/whl/cu128 \
  --extra-index-url https://pypi.org/simple \
  -r logs/venv-requirements.txt
conda run -n moss-awq3060 pip install --force-reinstall --no-deps \
  --index-url https://pypi.org/simple torchcodec==0.8.1
conda run -n moss-awq3060 pip install -e MOSS-TTS --no-deps --no-build-isolation
## 2. Model (already downloaded; to redownload):
conda run -n moss-awq3060 hf download lemuriandezapada/MOSS-TTS-v1.5-awq-int4 \
  --local-dir $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4
## 3. Env check / regression test / smoke / benchmark / register / comparison:
CUDA_VISIBLE_DEVICES=2 conda run -n moss-awq3060 python src/check_environment.py
CUDA_VISIBLE_DEVICES=2 conda run -n moss-awq3060 python tests/test_awq_forward.py
./run_single.sh
./run_benchmark.sh
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  conda run -n moss-awq3060 python src/test_register_awq.py
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  conda run -n moss-awq3060 python src/test_register_awq.py --seed 500 \
  --out benchmarks/register_awq_501
conda run -n moss-awq3060 python src/make_comparison.py

…[truncated 3714 chars]