The question
The HQ4 preview tried to save memory while protecting layers that mattered to audio. Text-token fidelity, codec behavior and staged-memory probes were checked separately, with calibration and physical-device limits still open.
From the original notebook
Background
Objective: highest-quality practical MOSS-TTS-v1.5 quant for a single RTX 3060 12 GB — INT4 only the Qwen3 transformer backbone (252 projection linears), everything audio-sensitive at BF16. Priorities: speech quality > similarity > pronunciation > prosody > stability > 12 GB fit > speed > size.
Source revisions: MOSS-TTS repo 934d682, model OpenMOSS-Team/MOSS-TTS-v1.5
(cdd3b91), codec OpenMOSS-Team/MOSS-Audio-Tokenizer (3cd226b).
Toolchain: torch 2.9.1+cu128, transformers 5.0.0, auto-round 0.15.1, Python 3.12.
Build GPU: RTX 3090 (no 3060 on host — acceptance via 0.5 memory fraction ≈ 12 GB).
Journey
- Spec + arch check — config verified exactly (Qwen3: 36 layers, h4096, 32 heads / 8 KV, vocab 155648; 32 audio embeddings, 33 LM heads, BF16). Inspection script confirmed 252 target linears = 81.8% of 8.49B params.
- Calibration — 512 packed multilingual TTS samples (EN 45 / ES 25 / ZH 10 / other 10 / dense 10), each tokenizer-verified ≥1024 tokens (AutoRound drops shorter samples; first attempt with ~100-char samples silently yielded zero data).
- Environment fights (documented, all resolved) —
flash-attn source build failed → SDPA fallback (spec-allowed);
torchcodec needed
nvidia-npp-cu12+ RTLD_GLOBAL preload (system CUDA shadowing); asymmetric W4A16 export crashes in auto-round 0.15.1 (vendoredWQLinear_GEMMrejects its ownweight_dtypekwarg) andawqhas no torch-2.9 wheel → symmetric fallback (spec-allowed). - AutoRound footguns — path-based loading wraps the backbone as
Qwen3ForCausalLM with a random lm_head (pass a loaded object instead);
language_config._name_or_path="Qwen/Qwen3-8B"triggered a 16 GB re-download on every construction (stub-dir fix); pointing the probe at the MOSS dir misclassifies as multimodal (processor_config.json) → LLM-calibrator fix. - Build saga — full-spec run (512 samples / 200 iters) managed ~1055 s/block (≈10 h ETA); a parallel fast run (128 samples / 100 iters / grad-accum 4, symmetric, 60 min) produced the artifact under test here: 6.3 GB, 2 shards, 252/252 INT4, protected modules BF16. Below spec minimums (256/200) → preview status until the full-spec build A/Bs against it.
- Sampled A/B — 33 prompts (EN/ES/ZH/mixed, clones, long-form to ~85s,
pause markers, duration control): 33/33 OK both models, median duration
delta 1.3s. Outliers to 27s (
longform-es-1,duration-short) are sampling chaos (temps 1.5/1.7), not signal — hence the greedy test below. - Greedy parity (this entry’s core test) — 10 prompts, temperatures 0,
token-sequence comparison. Text channel 96.5% (7/10 bit-identical over
full 1200-token runs); audio 21% (greedy cascade expected). Notable:
en-long-1HQ4 stops cleanly at 268 tokens / 18.7s while BF16 greedy runaways to the 1200 cap / 93.4s — greedy-mode EOS difference, sampled behavior matches (~20s both). - 12 GB envelope (this entry’s other test) — first attempt OOMed at 11.59 GB; root cause was decode-phase coexistence (6.3 GB generator + 6.7 GB FP32 codec), not load transient (direct CUDA load peaks = settled 6.24 GB). Fix: park generator on CPU during decode. Result: 3/3 PASS, 6.7–6.8 GiB peak. Detour recorded: CPU-first load selects slow pure-torch kernels (3.7× slower); direct CUDA load selects fast Triton kernels.
- Verdict — fast build is faithful (text 96.5%) and fits 12 GB with margin. Remaining: blind-listening gate, full-spec (256/200) build + A/B, then follow-ups in risk order (codec BF16 → head/embedding INT8 → codec INT8).
Open questions
- Blind-listening verdict on the paired WAVs (human ears required).
- 3060 RTF projection (~2–2.5) needs a real-3060 or bandwidth-throttled run to confirm.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.