Scope note: the original task targeted an RTX 3060 12GB, but the lab machine has 3× RTX 3090 and no 3060. Per user direction the work was retargeted to one free RTX 3090 (physical GPU 2). File names keep the
3060identifier for checklist compatibility; every measurement below is from the 3090 unless stated otherwise.
Hardware
- GPU: 1× NVIDIA GeForce RTX 3090 24GB (
CUDA_VISIBLE_DEVICES=2), idle at start - Driver 595.58.03, CUDA runtime 12.8, compute capability 8.6 (Ampere)
- Machine also holds GPU 0/1 (RTX 3090, untouched); OS Ubuntu, ffmpeg 6.1.1
Software versions
| Package | Version | Notes |
|---|---|---|
| Python | 3.11.16 (conda env moss-awq3060) |
conda create -n moss-awq3060 python=3.11 + pip stack |
| torch / torchaudio | 2.9.1+cu128 | PyTorch cu128 index, same as BF16 baseline |
| transformers | 5.17.0 | Baseline pins 5.0.0; 5.17 required by gptqmodel 7.5 (>=5.14) |
| gptqmodel | 7.5.0 | AWQ kernel backend (Marlin) |
| torchao | 0.16.0+cu128 | 0.18 breaks on torch 2.9 (ScalingType import) |
| accelerate | 1.15.0 | Same as baseline |
| safetensors / huggingface_hub | 0.8.0 / 1.32.0 | Same as baseline |
| numpy | 2.2.6 | Baseline 2.1.0; 2.2.6 required by gptqmodel 7.5 |
| torchcodec | 0.8.1 (PyPI) | +cu128 build fails to load libtorchcodec here (no system FFmpeg dev libs); PyPI build works |
| MOSS-TTS source | OpenMOSS/MOSS-TTS @ 934d682, editable install --no-deps |
Cloned into MOSS-TTS/ inside the lab |
Full freeze: logs/conda-pip-freeze.txt, logs/conda-list.txt.
pip check deltas vs baseline: only the known moss-tts notes (gradio
not installed, numpy/safetensors pins newer than pyproject).
Checkpoint
lemuriandezapada/MOSS-TTS-v1.5-awq-int4→$WORKSTATION/models/MOSS-TTS-v1.5-AWQ4- Single
model.safetensors, 6.3 GB on disk (vs 16 GB BF16, 6.3 GB HQ4) - Qwen3-8B backbone quantized with AutoAWQ 0.2.9 (GEMM, 4-bit, group 128,
zero-point); 33
lm_heads, 32emb_ext,embed_tokensuntouched (modules_to_not_convert); codec is the separate BF16MOSS-Audio-Tokenizer(read-only reused from$WORKSTATION/MOSS-TTS/weights/, 6.7 GB)
How AWQ was loaded
transformers-native AWQ path (AwqQuantizer, trust_remote_code=True,
attn_implementation="sdpa", device_map="cuda:0"), kernels from
gptqmodel 7.5.0, which auto-selects AwqMarlinLinear (Ampere Marlin;
JIT-compiled once to $WORKSTATION/.cache/gptqmodel, ~104 s first run, cached after).
Two load-time interventions were required (see Failures/Fixes):
- bf16 compute override. transformers 5.17 force-casts AWQ to fp16, but
Qwen3 outlier activations reach ~85504 (measured, decoder layer 7) and
overflow fp16 (max 65504) → inf/NaN logits → CUDA device-side assert in
sampling. gptqmodel Marlin supports bf16, so
run_awq_3060.pymonkeypatchesAwqQuantizer.update_dtypeto pass bf16 through. Model runs int4 weights / bf16 compute / bf16 untouched modules. - No AutoAWQ package. transformers ≥5 loads AWQ via gptqmodel, not
AutoAWQ; installing AutoAWQ 0.2.9 would have downgraded transformers to 4.x
and broken the v1.5 custom modeling code. Verified: 252 quantized linears
(
AwqMarlinLinear), 33 plainLinearheads, 33 fp→bf16 embeddings.
Staged-loading architecture
src/run_awq_3060.py::run_single implements generate → full unload → decode:
- Processor on CPU; AWQ transformer on
cuda:0(codec parked on CPU during generation; reference-encode stages the codec briefly for clone prompts). - Timed generation under
torch.inference_mode()(no autocast — native bf16), peak stats reset around it. - Token IDs → CPU;
delmodel + CUDA tensors;gc.collect();empty_cache();ipc_collect(); synchronize; residual recorded. - Codec →
cuda:0only afterwards; peak stats reset; timed decode → WAV.
Known wart: step 3 retains exactly one model’s VRAM (~6.24 GiB) due to an unidentified gptqmodel-side reference (see Failures). Decode peak therefore reads ~13.04 GiB (retained transformer + codec) instead of ~6.8 GiB codec-only. No unbounded accumulation (reload settles back to 6.24 GiB), and the 33-prompt benchmark uses a load-once design so this never compounds.
Single-prompt smoke result
- Text (Spanish): “Hola. Esta es una prueba de síntesis de voz en español. Quiero una pronunciación natural, clara y relajada.” (seed 1234)
- Backend Marlin/bf16: valid non-silent WAV (see
audio/awq3060-single.wav) - Representative run: 8.9–11.2 s audio from 145–174 tokens (Marlin bf16 is run-to-run nondeterministic on sm86 — see below), gen ~4.4–8.8 s, decode ~0.26 s, end-to-end RTF ~0.46–0.81, 20–36 tok/s
VRAM measurements
(Single-prompt staged run, Marlin/bf16.)
| Phase | Peak allocated | Peak reserved |
|---|---|---|
| Generation (transformer + KV) | ~6.31 GiB | ~6.38 GiB |
| Post-unload residual | ~6.24 GiB | ~6.32 GiB |
| Decode (retained transformer + codec) | ~13.04 GiB | ~13.10 GiB |
Codec-only estimate: 13.04 − 6.24 ≈ 6.8 GiB. For comparison, BF16 peaks at 22.6 GiB combined (15.9 gen + codec) and HQ4 at ~13.2 GiB combined.
33-prompt benchmark
- Harness:
src/benchmark_awq_3060.pyreuses the exact 33 prompts, seeds (1234), sampling and per-prompt token budgets fromMOSS-TTS/tools/benchmark_quality.py(read-only import; prompts verifiedlen == 33). Reference-asset paths resolved against$WORKSTATION/MOSS-TTS. - Design: load-once (one model load for all prompts; load time reported
separately, never in RTF). Per prompt: fresh seed-1234 RNG (same procedure
as baseline), timed bf16/SDPA generation, CPU-staged tokens, timed decode,
per-prompt JSON persisted immediately (
benchmarks/awq/). - RTF definitions:
rtf/pooled_rtfare generation-only to match the historical harness;end_to_end_rtfadds decode (the honest number).
Result: 33/33 ok, 0 failed, 0 silent. Full per-prompt records:
benchmarks/awq/benchmark_awq.json (+ results/benchmark_awq.json copy);
WAVs in benchmarks/awq/wav/.
| Metric (gen-only RTF, suite def.) | BF16 | HQ4 | AWQ (Marlin/bf16) |
|---|---|---|---|
| Median RTF | 0.59 | 0.96 | 0.47 |
| Pooled RTF (Σgen/Σaudio) | 0.59 | 0.88 | 0.43 |
| Tokens/s (pooled) | 25.5 | 17.0 | 35.8 |
| Audio seconds | 446.6 | 446.4 | 410.6 |
| Compute seconds (gen) | 262.5 | 394.5 | 174.5 (+3.8 decode) |
| Peak VRAM (suite max) | 22.77 GiB | 13.19 GiB | 6.49 gen / 13.05 decode-combined |
| Weights on disk | 16 GB | 6.3 GB | 6.3 GB |
Honest end-to-end numbers (AWQ only; historical harness never timed decode): median e2e RTF 0.48, pooled e2e RTF 0.43 (decode adds ~2% — 3.8 s over 410.6 s audio). Gen-time p50 4.07 s, p95 12.15 s; model load 4.0 s (once, excluded from RTF). Median audio 7.76 s.
Spanish quality observations
All 8 Spanish-bearing prompts produced valid non-silent audio
(es-conv-1, es-formal-1, es-question-1, es-numbers-1, es-long-1,
mix-en-es-1, clone-es-1, longform-es-1): RMS 0.034–0.173, no NaN/Inf,
sample rate 24 kHz throughout, ffprobe-clean. Durations track the baselines
within normal sampling spread (e.g. longform-es-1: AWQ 58.0 s vs BF16
85.0 s vs HQ4 57.9 s; es-long-1: 17.4 s vs 11.7 / 26.9 s) — no systematic
truncation (longform used 759 of 3000 budgeted tokens) and no silence or
token-leakage signature in any Spanish waveform. No transcript/WER evaluator
exists in the repos, so no WER claim is made; naturalness assessment beyond
these signal checks needs blind listening on the paired WAVs.
Voice/register-control findings
- Protocol mirrors
test_native_voice_hq4.py: instruction-guided female voice, EN+ES texts, seeds 1235+, F0 gate ≥195 Hz, ≤6 attempts/language, stop-on-pass; raw F0 via the same autocorrelation estimator. - Two seed sets mirroring the HQ4 protocol; reports in
benchmarks/register_awq/register_report.jsonandbenchmarks/register_awq_501/register_report.json.
| Suite | EN passes | ES passes | Combined |
|---|---|---|---|
| BF16 (seeds 1235) | 1/1 (226.4 Hz) | 0/3 (138–142 Hz) | 1/4 |
| HQ4 (1235 + 501) | 0/12 | 1/9 (196.7 @503) | 1/21 |
| AWQ (1235 + 501) | 1/9 (203.4 @1237) | 2/6 (198.3 @1237, 219.2 @503) | 3/15 |
AWQ restores usable instruction/register control where HQ4 had lost it (3× the passes in fewer attempts; female register reached in both languages), with a striking same-seed data point: seed 503/ES passes under both HQ4 (196.7 Hz, marginal) and AWQ (219.2 Hz, decisive). It does not fully match BF16’s first-try EN strength (226.4 Hz). Caveat: Marlin nondeterminism means a rerun draws different F0s; treat rates as single-draw estimates.
BF16 comparison
- Quality reference: BF16 holds the best register control (EN first-try 226 Hz) and is the naturalness anchor; AWQ Spanish signal checks are all healthy but blind listening is still needed for a naturalness verdict.
- Speed: AWQ-Marlin is ~1.4× faster at bs=1 decode (35.8 vs 25.5 tok/s; pooled RTF 0.43 vs 0.59) — measured, same suite/procedure/GPU class.
- VRAM: AWQ peaks at 13.05 GiB combined vs BF16’s 22.77 GiB (~43% less); transformer-only 6.49 GiB vs ~15.9 GiB.
- Disk: 6.3 GB vs 16 GB.
HQ4 comparison
- AWQ beats HQ4 on every measured axis: 2.1× faster (35.8 vs 17.0 tok/s; pooled RTF 0.43 vs 0.88), slightly lower peak VRAM (13.05 vs 13.19 GiB), same disk (6.3 GB), and materially better register control (3/15 vs 1/21 with both languages passing).
- HQ4’s AutoRound path is deterministic; AWQ-Marlin is not (see below). HQ4 remains a fallback only if exact reproducibility is required and RTF ~0.9 is acceptable.
Failures encountered
- No RTX 3060 on the machine (3× RTX 3090). Retargeted to one free 3090 per user direction; 3060-fit claims are projections, not measurements.
- fp16 activation overflow. transformers forces AWQ→fp16; Qwen3 outliers
hit ~85504 at decoder layer 7 → inf/NaN logits → CUDA device-side assert
in
torch.multinomialat generate step 0. Localized with a forward-pass NaN probe (inf-first at layer 7, clean embeds). - Marlin bf16 nondeterminism on sm86. Two identical in-process forwards differ by up to 23.0 in logits; same-seed generations diverge run to run (e.g. 174 vs 158 tokens). gptqmodel warns bf16-Marlin is a pre-SM90 compatibility path.
- Deterministic-backend bake-off failed on speed/correctness. TORCH_AWQ: deterministic, 1.9 tok/s (RTF 8.8). GEMM/bf16: deterministic, 3.1 tok/s (RTF 5.2). EXLLAMA_V2: deterministic but NaN logits. None viable as primary; Marlin kept as primary with documented caveat.
- Staged-unload retention. Post-generation
del+ gc + empty_cache retains exactly one model’s VRAM (6.24 GiB), backend-independent; decode peak reads ~13.04 GiB instead of ~6.8 GiB. No unbounded growth (reloads settle back to 6.24 GiB). Root holder unidentified after gc census, weakref census, referrer BFS, thread-stack walk, CUDA allocation snapshot, and inspection of custom modeling code, accelerate state, and gptqmodel post_init (no global caches found). Suspected invisible-to-gc reference in the quantizer load path. This is a 3060 blocker: per-item staged inference would peak ~19.5 GiB (stale + fresh + codec), exceeding 12 GB. - Dependency churn.
pip install gptqmodelsilently upgraded transformers 5.0→5.17 and numpy 2.1→2.2.6 (detected via before/after freeze; kept deliberately — gptqmodel 7.5 requires transformers ≥5.14 and older gptqmodel has no usable wheel here). torchao 0.18 broke on torch 2.9 (ScalingType); pinned 0.16.0. torchcodec+cu128couldn’t load libtorchcodec (no system FFmpeg libs); PyPI 0.8.1 works. - venv → conda migration mid-task per user request; stack replicated
pin-for-pin (106 packages) into
moss-awq3060, smoke-verified.
Fixes applied
AwqQuantizer.update_dtypemonkeypatch (bf16 passthrough) + bf16 model load — the fp16-overflow fix, locked bytests/test_awq_forward.py.- SDPA forced,
torch.inference_mode(), CUDA sync around all timings, peak-stat resets per phase, CPU token staging, immediate per-prompt JSON persistence. - Load-once benchmark (amortizes the 4 s load; immune to the retention quirk); per-prompt RNG reseeded to 1234 exactly like the baseline.
--backendswitch for kernel experiments (auto/marlin/torch_awq/ exllama_v2/gemm/gemm_triton/torch).
Recommended production configuration
AWQ INT4 + Marlin/bf16 + SDPA, load-once, batch 1, on a 24GB card — faster than BF16, far less VRAM, better register control than HQ4:
conda run -n moss-awq3060 python src/benchmark_awq_3060.py # or run_single.sh
Caveats: (a) nondeterministic outputs run-to-run — fine for serving, noisy
for eval; (b) keep transformers==5.17.0 + gptqmodel==7.5.0 pinned as in
logs/conda-pip-freeze.txt; (c) do NOT attempt 12GB deployment until the
unload-retention issue is resolved (needs ≤6.8 GiB codec-only decode;
currently ~13 GiB). Next optimization to try: Marlin fp16 with outlier
taming (unlikely — outliers are structural), or exllamav3/qqq AWQ backends
in newer gptqmodel for a deterministic fast path.
Exact reproduction commands
## 0. Prereqs: RTX 3090 GPU 2 free; miniforge3 conda; ffmpeg installed.
## 1. Env (already created; to recreate):
conda create -y -n moss-awq3060 python=3.11 pip
conda run -n moss-awq3060 pip install \
--index-url https://download.pytorch.org/whl/cu128 \
--extra-index-url https://pypi.org/simple \
-r logs/venv-requirements.txt
conda run -n moss-awq3060 pip install --force-reinstall --no-deps \
--index-url https://pypi.org/simple torchcodec==0.8.1
conda run -n moss-awq3060 pip install -e MOSS-TTS --no-deps --no-build-isolation
## 2. Model (already downloaded; to redownload):
conda run -n moss-awq3060 hf download lemuriandezapada/MOSS-TTS-v1.5-awq-int4 \
--local-dir $WORKSTATION/models/MOSS-TTS-v1.5-AWQ4
## 3. Env check / regression test / smoke / benchmark / register / comparison:
CUDA_VISIBLE_DEVICES=2 conda run -n moss-awq3060 python src/check_environment.py
CUDA_VISIBLE_DEVICES=2 conda run -n moss-awq3060 python tests/test_awq_forward.py
./run_single.sh
./run_benchmark.sh
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
conda run -n moss-awq3060 python src/test_register_awq.py
CUDA_VISIBLE_DEVICES=2 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
conda run -n moss-awq3060 python src/test_register_awq.py --seed 500 \
--out benchmarks/register_awq_501
conda run -n moss-awq3060 python src/make_comparison.py
…[truncated 3714 chars]