Journal

A draft head, a rollback fix, and a midnight incident

The rollback fix brought draft decoding to parity at temperature 0.7 and a 24% gain at greedy sampling.

Vlad / experimentos.
The boundary

The scheduled 200K validation was invalidated by an external disk reorganization. It did not pass.

The question

Draft decoding exposed a rollback issue, and the fix needed a fresh comparison at different sampling settings. The later disk incident is retained because it invalidated the scheduled long-context validation.

From the original notebook

Plain-language summary: FINDINGS.md — what won, why MTP did not ship, what broke at midnight.

  • Date: 2026-09-07 → 2026-09-08 (export + A/B + scheduled validation run)
  • Host: 12700 — 3x RTX 3090 (72 GB aggregate VRAM), 20 CPUs, Ubuntu Linux
  • Target model: orcarouter/Qwen3.8-Flash-Next-Uncensored-FP8 (abliterated 125B-A6B MoE) → production quant AD-4.27bpw Q4_K_M-M64-agentic-v2 (see sibling experiment)
  • Draft model: MTP (NextN) head exported from the same checkpoint → mtp-Qwen38-Unc-Q8_0.gguf (4.1 GB, 34 tensors) and mtp-Qwen38-Unc-Q4_K_M.gguf (2.78 GB)
  • Inference stack: llama.cpp flash-next fork at WORKSTATION/llama.cpp-flash-next (master 9723942 + local uncommitted patch), --repack --op-offload, -sm layer -ts 1,1,1
  • Router: qwen36-multi-proxy on :12434, family qwen38-flash-next-unc-q4km-agentic-v2, backend :12702

Two questions answered: (1) can the Qwen3.8 “qwen4exp” MTP/NextN draft head be exported to GGUF and made to run in llama.cpp, and (2) does draft-MTP speculative decoding beat plain decode on this 3x3090 stack at production sampling. A third section documents what happened when the 200K-context validation was scheduled to run unattended at midnight.

1. Draft-head export (upstream PR #27836 + local fixes)

Upstream PR ggml-org/llama.cpp#27836 (“qwen4exp: add NextN/MTP draft head”) applied via git apply, plus fixes the PR was missing (all working-tree only, never committed):

File Change
conversion/qwen4exp.py Dropped supports_mtp_export=False; fused mtp.fc_embedding + mtp.fc_hidden → eh_proj [5120,2560] (mathematically identical W_e@e + W_h@h == [W_e|W_h]@concat); HC mixer tensor remap; GGUF param padding for dense MTP
gguf-py (constants, tensor_mapping) NEXTN_HC_HEAD_{NORM,DOWN,UP} tensor names
src/llama-arch.{h,cpp}, llama-model.{h,cpp} Tensor infos; nextn layer gains hc_head_*; QWEN4EXP added to mtp_on_hybrid_qwen (draft ctx → plain KV, not hybrid)
src/models/qwen4exp.cpp PR body + required local fix: load_arch_tensors probes blk.0.hc_attn_norm.weight; when absent (draft-only GGUF) trunk hc_head_* become optional — else draft load dies on output_hc_norm.weight (hit and fixed during smoke)

Pipeline: selective 30-shard/32 GB download → convert_hf_to_gguf.py --mtp --outtype q8_0 (exit 0, 4.12 GiB; MTP-head tensors are BF16 in the FP8 repo, so no dequant-order hazard for the fused eh_proj) → llama-quantize --allow-requantize → Q4_K_M 2.78 GB. Smoke at ctx 32768: draft acceptance 0.615 (64/104), mean accepted len 2.83, coherent reply — draft genuinely active.

GGUF facts: arch qwen4exp, block_count 49, nextn_predict_layers=1, 512 experts, n_embd 2560; 34 tensors = token_embd + output + one full-attn + 512-expert MoE block under blk.48 + nextn.*.

2. Track 1: ported upstream #28123 (conv-state rollback) — the A/B that mattered

The first smoke ran at 21 t/s (2.7x slower than baseline). Root cause: recurrent-state rollback was missing for qwen4exp — every draft rejection discarded the whole sequence. Ported merged upstream #28123 (llm_arch_supports_rs_rollback + slotted conv-state store, the fork’s existing TAG_RECURRENT_ROLLBACK_SPLITS pattern). No-op when n_rs_seq=0; production no-draft path provably untouched (n_layer_nextn=0 → flags=0 everywhere).

A/B on 3x3090, ctx 32768, taskset-pinned, draft on CUDA1, median-of-5 (scripts/ab_*.sh, data/ab_*.json):

config temp 0.7 (median) temp 0 (greedy)
baseline no-draft 54.3 t/s 57.4 t/s
draft-MTP n=3 + rollback 55.3 t/s — parity 71.0 t/s (+24%)
draft-MTP n=6 + rollback 47–52 t/s (worse) 59.6 t/s (worse)
  • Rollback proven engaged (“bounded partial sequence removal” at startup); stability 5/5 identical-prompt chats with prompt-cache restore — the exact path that aborts with on-device checkpoints (#28118 approach, rejected).
  • Acceptance ~0.53–0.62 at both temps (upstream reports 0.79 at p-min 0.7 on the base model) — the Unc fine-tune / base-head mismatch caps gains.
  • Box-noise lesson: temp-0.7 runs swing ±20% under CPU contention from other sessions; temp-0 runs are rock-stable (<1% spread). Medians mandatory; never compare across core sets.
  • Earlier no-rollback comparison for reference (data/bench_v4_65k*.json): v4+MTP lost 20–48% vs v3 at ctx 65536 despite 60–75% acceptance.

3. Decision: keep v3 (no draft) on :12702

At the production sampling preset (temp 0.7) MTP is parity — no rollout on speed alone. At temp 0 (greedy/code/agentic) n=3 gives a stable +24% → rollout candidate only if greedy workloads matter. Both user presets are acceptance-hostile anyway: thinking runs at temp 1.0 (maximum entropy), and instruct’s presence 1.5 distorts the target distribution away from draftable argmax paths. data/DECIDE.json ranking (warmed mean gen t/s): v3_65k 56.9 > v3_200k 55.6 (shipped) > mtp_v4_65k 39.2 > other MTP variants 35.8–38.5.

4. The 200K validation that never ran — midnight incident post-mortem

out/PLAN_CTX200K.md specified a 4-phase validation (fit ladder 204800→65536 with a 1.5 GiB/GPU floor, speed A/B under the two user presets, cyber suite, 198K qualification) with trap-guaranteed production restore, scheduled one-shot via cron 0 0 8 9 * (2026-09-08 00:00 NZST). The cron fired at 00:00:01 exactly; the failure was environmental, and the runner caught it correctly:

Time Event Evidence
~19:45 Production :12702 stopped externally (SIGTERM, graceful exit) v3_sharp.log ends in cleaning up before exit
~19:48 Third party reorganized WORKSTATION/models: 33 shards consolidated into one 94.5 GB file; deleted the Q4_K_M draft, the download cache, and the UD-IQ4_XS tree (~110 GB freed) mtimes; draft ls fails; df 141→251 GB free
20:20 Another session started a manual proxy on :12434 (11 models, no allowlist) → systemd unit crash-loops on “port already in use” every 10 s pid 1726409 vs systemd journal
00:00:01 Cron fires; preflight correctly aborts: “sharp template or draft missing” (draft deleted) ctx200k_run.log
00:00:03 Trap-restore attempts v3 ×2; both die instantly: the script’s MODEL_PATH glob finds 2 candidates (mmproj + consolidated single file) → die; blind 15-min health waits ×2 ctx200k_v3restored.log
00:30 Runner declares PRODUCTION DOWN + Telegram alert (correct given what it could see) ctx200k_run.log

Post-mortem findings (runner hardening lessons):

  1. Glob-based MODEL_PATH is fragile — a directory with 2 GGUF candidates (weights + mmproj) must resolve to the entrypoint explicitly, else fail fast, not after 30 min of blind waiting.
  2. Health waits should poll with early-exit on dead child processes, not sleep the full window.
  3. Preflight was the hero: validating artifacts before stopping production is what prevented a 2.5 h wasted downtime window — but the draft-file check ran before checking production was already down; ordering preflight checks by blast radius matters.
  4. Duplicate log lines (tee + cron redirect) are cosmetic but double every diagnostic.
  5. Shared-box reality: another actor’s disk cleanup can invalidate a scheduled job’s inputs hours before it fires. Scheduled jobs should re-verify inputs and re-check assumptions at fire time.

Resolution: production was restored externally the next day — :12702 healthy on the consolidated single-file GGUF + mmproj, with the sharp chat template (--chat-template-file), pid verified 2026-09-09. The Q4_K_M draft is gone (Q8_0 draft survives, 3.9 GB; speed-equivalent per DECIDE.json); re-export would need the 32 GB selective download again. The cron line remains (fires next on 2027-09-08; remove for hygiene).

5. Router deployment notes (same arc, worth keeping)

  • Family-only exposure works via env on the systemd unit: MODEL_FAMILY_ALLOWLIST + DISABLED_AUXILIARY_MODEL_IDS — but the drop-in must sort last (zz-*.conf) because the stock laguna-env.conf loads /etc/laguna-s21-proxy.env afterwards and overwrites DISABLED_AUXILIARY_MODEL_IDS.
  • The family JSON uses llama.cpp’s key name repeat_penalty (not repetition_penalty).
  • Sharp template on the backend: CHAT_TEMPLATE_FILE=WORKSTATION/chat_template.jinja through the v3 launch script — required because the baked template ignores enable_thinking on this abliterated model (template finding #1 of qwen-next-flash).
  • Presets deployed (verified live via /v1/presets): thinking 1.0/0.95/20/min_p 0/pres 0/rep 1.0; instruct 0.7/0.80/20/min_p 0/pres 1.5/rep 1.0.

6. Honest caveats

  • All MTP numbers are ctx 32768 single-node medians; production 200K+MTP fit was never measured (prior attempt OOMs, and the validation that would have settled it never ran).
  • The draft head v1 attends densely in the MTP block (trunk QSA-indexer pruning not wired into graph_mtp); numerically a superset, matters only for long-ctx draft cost.
  • Q4_K_M draft came from q8_0 requant (--allow-requantize double quant); from-BF16 would be marginally better.
  • Acceptance capped ~0.6 suggests the draft head mismatches the abliterated fine-tune; p-min tuning untested.
  • Patches are uncommitted in WORKSTATION/llama.cpp-flash-next by instruction; the output_hc_norm mtp_only fix is worth upstreaming to #27836.

Reproduction

  1. Export: apply #27836 + §1 fixes, then selective download → convert_hf_to_gguf.py --mtp --outtype q8_0 → llama-quantize --allow-requantize (job log: quant-jobs/mtp-unc-export/logs/).
  2. Smoke: scripts/smoke_draft_mtp.sh (target + draft, port 12788).
  3. A/B: scripts/ab_baseline.sh / scripts/ab_draft.sh (taskset-pinned, median-of-5).
  4. Scheduled validation (as designed): scripts/cron_wrapper_ctx200k.sh → scripts/run_ctx200k_validation.sh; probes scripts/speed_ab_200k.py, scripts/qual_200k.py, suite scripts/bench_ctx200k_mtp.py.
  5. Data: data/ab_*.json (Track 1), data/decide_*.json + data/DECIDE.json (ranking), data/bench_v3*.json / data/bench_v4*.json (no-rollback comparison).

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS