Back to the experimentSupporting notebook

HQ4 PR #1 hardening — workstation validation

Source: moss-tts-hq4/pr1-hardening-validation-2026-09-18/README.md · revision 6550ead3945b

Date: 2026-09-18. Repo: groxaxo/moss-tts-hq4-quant, branch fix/release-reproducibility-hardening-v2 at 7982ede (7982edee4f10b293db8a7e4d7293661a69679692), base main at bfd27b1. PR #1 (draft): “Harden HQ4 release provenance, validation, and 12GB staging claims”. Workstation checkout: WORKSTATION/moss-tts-hq4-quant.

Question answered: do the PR’s “local GPU gates before merge” pass on the Ubuntu workstation — py_compile, bash -n, verify_quantization_map.py, and the staged 12 GB inference path?

Measured result: all listed gates pass. Verifier: 252/252 INT4 projections, 211 protected BF16 params, bits=4 / g=128 / sym=true. Staged text-only run under a 0.5 memory fraction on one RTX 3090 peaks at 6.63 GiB (vs 13.01 GiB pre-staging), with the generator fully parked on CPU (0.01 GiB) during codec decode. No physical RTX 3060 exists on this host, so the true 3060 acceptance run remains open by design.

Workflow

flowchart LR
    A[Locate repo +<br/>fetch PR branch] --> B[Checkout 7982ede<br/>review 12-file diff]
    B --> C[Gate 1: py_compile<br/>4 tools]
    C --> D[Gate 2: bash -n<br/>4 scripts]
    D --> E[Gate 3: verifier<br/>GPU 2, real artifact]
    E --> F[Gate 4: staged run<br/>memfrac 0.5, GPU 1]
    F --> G[Phase-VRAM probe<br/>run_single + rec]
    G --> H[Quantizer probes<br/>refusal + recipe]
    H --> I[This record<br/>+ PR update]

Step outcomes (observed, in order):

# Action Evidence Outcome
1 Fetched origin, checked out PR tip git log, diffstat 12 files +268/-138 On 7982ede, clean tree
2 py_compile on 4 touched tools Console PY_COMPILE_OK Pass
3 bash -n on 4 ops scripts Console BASH_N_OK x4 Pass
4 Verifier on GPU 2 vs real artifact verifier-tail VERIFY PASSED, exit 0
5 Staged run, memfrac 0.5, GPU 1 sample, JSON Exit 0, peak 6.647 GiB
6 Phase-VRAM capture via run_single(rec) sample, phases below Exit 0, peak 6.634 GiB
7 Quantizer refusal/recipe probes (CPU-only) Console *_OK x5 All pass
8 Voice-test flags + wav sanity --help, peak 0.79, no NaN, not silent Pass

What was tested

Environment: 3× RTX 3090 24 GB (driver 595.58.03), Python 3.12.3, torch 2.9.1+cu128, transformers 5.0.0, auto-round 0.15.1 (see versions). Runtime venv, weights, and artifact live under WORKSTATION/MOSS-TTS — moss-tts-hq4-quant ships code only, so the GPU gates ran the PR-branch scripts with the MOSS-TTS venv and absolute model paths.

Gate Command (equivalent) Result
py_compile python3 -m py_compile tools/{infer_hybrid_3060,quantize_moss_autoround_awq,test_native_voice_hq4,verify_quantization_map}.py Pass
bash -n hq4_chain_gpu2.sh, monitor-hq4.sh, both retry-wrapper-hq4-fast.sh Pass x4
Verifier CUDA_VISIBLE_DEVICES=2 … verify_quantization_map.py --model artifacts/… VERIFY PASSED, exit 0
Staged 12 GB infer_hybrid_3060.py … --memory-fraction 0.5 --seed 1234 (text-only) Exit 0, 5.36 s audio, RTF 1.76

Verifier detail: 36 layers → 252/252 quantized projections, 252/252 qweight tensors, 211 protected BF16 parameters, quant_method=auto-round.

Phase VRAM (GiB, allocated / peak), second run:

Phase Allocated Peak
load 6.237 6.237
prefill_ready 6.237 6.237
generate 6.250 6.590
generator_parked_cpu 0.013 6.590
decode 6.625 6.634

Overall peak 6.634 GiB — the pre-staging benchmark peaked at 13.01 GiB with generator+codec coexisting at decode (see benchmark_results_hq4.json in the artifact dir). The CPU-parking fix recovers ~6.4 GiB at decode.

Quantizer probes (CPU-only, throwaway dirs under /tmp, no GPU, no quant run): namespace remap, missing-calibration FileNotFoundError pointing at tools/build_moss_calibration.py, untracked-export refusal, stale-export fingerprint-mismatch refusal, and recipe contents (cal SHA256, W4A16/sym settings, dep versions) — all pass.

Runs and timing

Run GPU Result Wall
Verifier vs MOSS-TTS-v1.5-HQ4-AWQ 2 (as cuda:0) VERIFY PASSED ~40 s incl. load
Staged envelope (staged-12gb-envelope.wav) 1 (as cuda:0), memfrac 0.5 Exit 0, 101 tokens 19.2 s total, 9.4 s gen
Phase capture (phase-capture-run.wav) 1 (as cuda:0), memfrac 0.5 Exit 0 ~20 s

Limits and reproduction boundary

No merge or deployment performed; PR #1 left draft and open.