Back to the experimentSupporting notebook

Unc Q4_K_M MTP draft head: export, speculative-rollback A/B, and the 200K validation incident

Source: qwen38/flash-next-unc-mtp-draft-2026-09-07/README.md · revision 6550ead3945b

Plain-language summary: FINDINGS.md — what won, why MTP did not ship, what broke at midnight.

Two questions answered: (1) can the Qwen3.8 “qwen4exp” MTP/NextN draft head be exported to GGUF and made to run in llama.cpp, and (2) does draft-MTP speculative decoding beat plain decode on this 3x3090 stack at production sampling. A third section documents what happened when the 200K-context validation was scheduled to run unattended at midnight.

1. Draft-head export (upstream PR #27836 + local fixes)

Upstream PR ggml-org/llama.cpp#27836 (“qwen4exp: add NextN/MTP draft head”) applied via git apply, plus fixes the PR was missing (all working-tree only, never committed):

File Change
conversion/qwen4exp.py Dropped supports_mtp_export=False; fused mtp.fc_embedding + mtp.fc_hidden → eh_proj [5120,2560] (mathematically identical W_e@e + W_h@h == [W_e|W_h]@concat); HC mixer tensor remap; GGUF param padding for dense MTP
gguf-py (constants, tensor_mapping) NEXTN_HC_HEAD_{NORM,DOWN,UP} tensor names
src/llama-arch.{h,cpp}, llama-model.{h,cpp} Tensor infos; nextn layer gains hc_head_*; QWEN4EXP added to mtp_on_hybrid_qwen (draft ctx → plain KV, not hybrid)
src/models/qwen4exp.cpp PR body + required local fix: load_arch_tensors probes blk.0.hc_attn_norm.weight; when absent (draft-only GGUF) trunk hc_head_* become optional — else draft load dies on output_hc_norm.weight (hit and fixed during smoke)

Pipeline: selective 30-shard/32 GB download → convert_hf_to_gguf.py --mtp --outtype q8_0 (exit 0, 4.12 GiB; MTP-head tensors are BF16 in the FP8 repo, so no dequant-order hazard for the fused eh_proj) → llama-quantize --allow-requantize → Q4_K_M 2.78 GB. Smoke at ctx 32768: draft acceptance 0.615 (64/104), mean accepted len 2.83, coherent reply — draft genuinely active.

GGUF facts: arch qwen4exp, block_count 49, nextn_predict_layers=1, 512 experts, n_embd 2560; 34 tensors = token_embd + output + one full-attn + 512-expert MoE block under blk.48 + nextn.*.

2. Track 1: ported upstream #28123 (conv-state rollback) — the A/B that mattered

The first smoke ran at 21 t/s (2.7x slower than baseline). Root cause: recurrent-state rollback was missing for qwen4exp — every draft rejection discarded the whole sequence. Ported merged upstream #28123 (llm_arch_supports_rs_rollback + slotted conv-state store, the fork’s existing TAG_RECURRENT_ROLLBACK_SPLITS pattern). No-op when n_rs_seq=0; production no-draft path provably untouched (n_layer_nextn=0 → flags=0 everywhere).

A/B on 3x3090, ctx 32768, taskset-pinned, draft on CUDA1, median-of-5 (scripts/ab_*.sh, data/ab_*.json):

config temp 0.7 (median) temp 0 (greedy)
baseline no-draft 54.3 t/s 57.4 t/s
draft-MTP n=3 + rollback 55.3 t/s — parity 71.0 t/s (+24%)
draft-MTP n=6 + rollback 47–52 t/s (worse) 59.6 t/s (worse)

3. Decision: keep v3 (no draft) on :12702

At the production sampling preset (temp 0.7) MTP is parity — no rollout on speed alone. At temp 0 (greedy/code/agentic) n=3 gives a stable +24% → rollout candidate only if greedy workloads matter. Both user presets are acceptance-hostile anyway: thinking runs at temp 1.0 (maximum entropy), and instruct’s presence 1.5 distorts the target distribution away from draftable argmax paths. data/DECIDE.json ranking (warmed mean gen t/s): v3_65k 56.9 > v3_200k 55.6 (shipped) > mtp_v4_65k 39.2 > other MTP variants 35.8–38.5.

4. The 200K validation that never ran — midnight incident post-mortem

out/PLAN_CTX200K.md specified a 4-phase validation (fit ladder 204800→65536 with a 1.5 GiB/GPU floor, speed A/B under the two user presets, cyber suite, 198K qualification) with trap-guaranteed production restore, scheduled one-shot via cron 0 0 8 9 * (2026-09-08 00:00 NZST). The cron fired at 00:00:01 exactly; the failure was environmental, and the runner caught it correctly:

Time Event Evidence
~19:45 Production :12702 stopped externally (SIGTERM, graceful exit) v3_sharp.log ends in cleaning up before exit
~19:48 Third party reorganized WORKSTATION/models: 33 shards consolidated into one 94.5 GB file; deleted the Q4_K_M draft, the download cache, and the UD-IQ4_XS tree (~110 GB freed) mtimes; draft ls fails; df 141→251 GB free
20:20 Another session started a manual proxy on :12434 (11 models, no allowlist) → systemd unit crash-loops on “port already in use” every 10 s pid 1726409 vs systemd journal
00:00:01 Cron fires; preflight correctly aborts: “sharp template or draft missing” (draft deleted) ctx200k_run.log
00:00:03 Trap-restore attempts v3 ×2; both die instantly: the script’s MODEL_PATH glob finds 2 candidates (mmproj + consolidated single file) → die; blind 15-min health waits ×2 ctx200k_v3restored.log
00:30 Runner declares PRODUCTION DOWN + Telegram alert (correct given what it could see) ctx200k_run.log

Post-mortem findings (runner hardening lessons):

  1. Glob-based MODEL_PATH is fragile — a directory with 2 GGUF candidates (weights + mmproj) must resolve to the entrypoint explicitly, else fail fast, not after 30 min of blind waiting.
  2. Health waits should poll with early-exit on dead child processes, not sleep the full window.
  3. Preflight was the hero: validating artifacts before stopping production is what prevented a 2.5 h wasted downtime window — but the draft-file check ran before checking production was already down; ordering preflight checks by blast radius matters.
  4. Duplicate log lines (tee + cron redirect) are cosmetic but double every diagnostic.
  5. Shared-box reality: another actor’s disk cleanup can invalidate a scheduled job’s inputs hours before it fires. Scheduled jobs should re-verify inputs and re-check assumptions at fire time.

Resolution: production was restored externally the next day — :12702 healthy on the consolidated single-file GGUF + mmproj, with the sharp chat template (--chat-template-file), pid verified 2026-09-09. The Q4_K_M draft is gone (Q8_0 draft survives, 3.9 GB; speed-equivalent per DECIDE.json); re-export would need the 32 GB selective download again. The cron line remains (fires next on 2027-09-08; remove for hygiene).

5. Router deployment notes (same arc, worth keeping)

6. Honest caveats

Reproduction

  1. Export: apply #27836 + §1 fixes, then selective download → convert_hf_to_gguf.py --mtp --outtype q8_0 → llama-quantize --allow-requantize (job log: quant-jobs/mtp-unc-export/logs/).
  2. Smoke: scripts/smoke_draft_mtp.sh (target + draft, port 12788).
  3. A/B: scripts/ab_baseline.sh / scripts/ab_draft.sh (taskset-pinned, median-of-5).
  4. Scheduled validation (as designed): scripts/cron_wrapper_ctx200k.sh → scripts/run_ctx200k_validation.sh; probes scripts/speed_ab_200k.py, scripts/qual_200k.py, suite scripts/bench_ctx200k_mtp.py.
  5. Data: data/ab_*.json (Track 1), data/decide_*.json + data/DECIDE.json (ranking), data/bench_v3*.json / data/bench_v4*.json (no-rollback comparison).