Journal

Keeping the audio-sensitive layers intact

The HQ4 preview matched 96.5% of text tokens in the greedy test and passed three staged-memory probes at 6.7–6.8 GiB.

Vlad / experimentos.
The boundary

The fast build was below the full calibration specification. The 12GB envelope was tested on a 3090, not a physical 3060.

The question

The HQ4 preview tried to save memory while protecting layers that mattered to audio. Text-token fidelity, codec behavior and staged-memory probes were checked separately, with calibration and physical-device limits still open.

From the original notebook

Background

Objective: highest-quality practical MOSS-TTS-v1.5 quant for a single RTX 3060 12 GB — INT4 only the Qwen3 transformer backbone (252 projection linears), everything audio-sensitive at BF16. Priorities: speech quality > similarity > pronunciation > prosody > stability > 12 GB fit > speed > size.

Source revisions: MOSS-TTS repo 934d682, model OpenMOSS-Team/MOSS-TTS-v1.5 (cdd3b91), codec OpenMOSS-Team/MOSS-Audio-Tokenizer (3cd226b). Toolchain: torch 2.9.1+cu128, transformers 5.0.0, auto-round 0.15.1, Python 3.12. Build GPU: RTX 3090 (no 3060 on host — acceptance via 0.5 memory fraction ≈ 12 GB).

Journey

  1. Spec + arch check — config verified exactly (Qwen3: 36 layers, h4096, 32 heads / 8 KV, vocab 155648; 32 audio embeddings, 33 LM heads, BF16). Inspection script confirmed 252 target linears = 81.8% of 8.49B params.
  2. Calibration — 512 packed multilingual TTS samples (EN 45 / ES 25 / ZH 10 / other 10 / dense 10), each tokenizer-verified ≥1024 tokens (AutoRound drops shorter samples; first attempt with ~100-char samples silently yielded zero data).
  3. Environment fights (documented, all resolved) — flash-attn source build failed → SDPA fallback (spec-allowed); torchcodec needed nvidia-npp-cu12 + RTLD_GLOBAL preload (system CUDA shadowing); asymmetric W4A16 export crashes in auto-round 0.15.1 (vendored WQLinear_GEMM rejects its own weight_dtype kwarg) and awq has no torch-2.9 wheel → symmetric fallback (spec-allowed).
  4. AutoRound footguns — path-based loading wraps the backbone as Qwen3ForCausalLM with a random lm_head (pass a loaded object instead); language_config._name_or_path="Qwen/Qwen3-8B" triggered a 16 GB re-download on every construction (stub-dir fix); pointing the probe at the MOSS dir misclassifies as multimodal (processor_config.json) → LLM-calibrator fix.
  5. Build saga — full-spec run (512 samples / 200 iters) managed ~1055 s/block (≈10 h ETA); a parallel fast run (128 samples / 100 iters / grad-accum 4, symmetric, 60 min) produced the artifact under test here: 6.3 GB, 2 shards, 252/252 INT4, protected modules BF16. Below spec minimums (256/200) → preview status until the full-spec build A/Bs against it.
  6. Sampled A/B — 33 prompts (EN/ES/ZH/mixed, clones, long-form to ~85s, pause markers, duration control): 33/33 OK both models, median duration delta 1.3s. Outliers to 27s (longform-es-1, duration-short) are sampling chaos (temps 1.5/1.7), not signal — hence the greedy test below.
  7. Greedy parity (this entry’s core test) — 10 prompts, temperatures 0, token-sequence comparison. Text channel 96.5% (7/10 bit-identical over full 1200-token runs); audio 21% (greedy cascade expected). Notable: en-long-1 HQ4 stops cleanly at 268 tokens / 18.7s while BF16 greedy runaways to the 1200 cap / 93.4s — greedy-mode EOS difference, sampled behavior matches (~20s both).
  8. 12 GB envelope (this entry’s other test) — first attempt OOMed at 11.59 GB; root cause was decode-phase coexistence (6.3 GB generator + 6.7 GB FP32 codec), not load transient (direct CUDA load peaks = settled 6.24 GB). Fix: park generator on CPU during decode. Result: 3/3 PASS, 6.7–6.8 GiB peak. Detour recorded: CPU-first load selects slow pure-torch kernels (3.7× slower); direct CUDA load selects fast Triton kernels.
  9. Verdict — fast build is faithful (text 96.5%) and fits 12 GB with margin. Remaining: blind-listening gate, full-spec (256/200) build + A/B, then follow-ups in risk order (codec BF16 → head/embedding INT8 → codec INT8).

Open questions

  • Blind-listening verdict on the paired WAVs (human ears required).
  • 3060 RTF projection (~2–2.5) needs a real-3060 or bandwidth-throttled run to confirm.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS