The question
Draft decoding exposed a rollback issue, and the fix needed a fresh comparison at different sampling settings. The later disk incident is retained because it invalidated the scheduled long-context validation.
From the original notebook
Plain-language summary: FINDINGS.md — what won, why MTP did not ship, what broke at midnight.
- Date: 2026-09-07 → 2026-09-08 (export + A/B + scheduled validation run)
- Host:
12700— 3x RTX 3090 (72 GB aggregate VRAM), 20 CPUs, Ubuntu Linux - Target model:
orcarouter/Qwen3.8-Flash-Next-Uncensored-FP8(abliterated 125B-A6B MoE) → production quant AD-4.27bpw Q4_K_M-M64-agentic-v2 (see sibling experiment) - Draft model: MTP (NextN) head exported from the same checkpoint →
mtp-Qwen38-Unc-Q8_0.gguf(4.1 GB, 34 tensors) andmtp-Qwen38-Unc-Q4_K_M.gguf(2.78 GB) - Inference stack: llama.cpp flash-next fork at
WORKSTATION/llama.cpp-flash-next(master9723942+ local uncommitted patch),--repack --op-offload,-sm layer -ts 1,1,1 - Router:
qwen36-multi-proxyon:12434, familyqwen38-flash-next-unc-q4km-agentic-v2, backend:12702
Two questions answered: (1) can the Qwen3.8 “qwen4exp” MTP/NextN draft head be exported to GGUF and made to run in llama.cpp, and (2) does draft-MTP speculative decoding beat plain decode on this 3x3090 stack at production sampling. A third section documents what happened when the 200K-context validation was scheduled to run unattended at midnight.
1. Draft-head export (upstream PR #27836 + local fixes)
Upstream PR ggml-org/llama.cpp#27836 (“qwen4exp: add NextN/MTP draft head”) applied via git apply,
plus fixes the PR was missing (all working-tree only, never committed):
| File | Change |
|---|---|
conversion/qwen4exp.py |
Dropped supports_mtp_export=False; fused mtp.fc_embedding + mtp.fc_hidden → eh_proj [5120,2560] (mathematically identical W_e@e + W_h@h == [W_e|W_h]@concat); HC mixer tensor remap; GGUF param padding for dense MTP |
gguf-py (constants, tensor_mapping) |
NEXTN_HC_HEAD_{NORM,DOWN,UP} tensor names |
src/llama-arch.{h,cpp}, llama-model.{h,cpp} |
Tensor infos; nextn layer gains hc_head_*; QWEN4EXP added to mtp_on_hybrid_qwen (draft ctx → plain KV, not hybrid) |
src/models/qwen4exp.cpp |
PR body + required local fix: load_arch_tensors probes blk.0.hc_attn_norm.weight; when absent (draft-only GGUF) trunk hc_head_* become optional — else draft load dies on output_hc_norm.weight (hit and fixed during smoke) |
Pipeline: selective 30-shard/32 GB download → convert_hf_to_gguf.py --mtp --outtype q8_0 (exit 0,
4.12 GiB; MTP-head tensors are BF16 in the FP8 repo, so no dequant-order hazard for the fused eh_proj)
→ llama-quantize --allow-requantize → Q4_K_M 2.78 GB. Smoke at ctx 32768: draft acceptance
0.615 (64/104), mean accepted len 2.83, coherent reply — draft genuinely active.
GGUF facts: arch qwen4exp, block_count 49, nextn_predict_layers=1, 512 experts, n_embd 2560;
34 tensors = token_embd + output + one full-attn + 512-expert MoE block under blk.48 + nextn.*.
2. Track 1: ported upstream #28123 (conv-state rollback) — the A/B that mattered
The first smoke ran at 21 t/s (2.7x slower than baseline). Root cause: recurrent-state rollback was
missing for qwen4exp — every draft rejection discarded the whole sequence. Ported merged upstream
#28123 (llm_arch_supports_rs_rollback + slotted conv-state store, the fork’s existing
TAG_RECURRENT_ROLLBACK_SPLITS pattern). No-op when n_rs_seq=0; production no-draft path provably
untouched (n_layer_nextn=0 → flags=0 everywhere).
A/B on 3x3090, ctx 32768, taskset-pinned, draft on CUDA1, median-of-5 (scripts/ab_*.sh,
data/ab_*.json):
| config | temp 0.7 (median) | temp 0 (greedy) |
|---|---|---|
| baseline no-draft | 54.3 t/s | 57.4 t/s |
| draft-MTP n=3 + rollback | 55.3 t/s — parity | 71.0 t/s (+24%) |
| draft-MTP n=6 + rollback | 47–52 t/s (worse) | 59.6 t/s (worse) |
- Rollback proven engaged (“bounded partial sequence removal” at startup); stability 5/5 identical-prompt chats with prompt-cache restore — the exact path that aborts with on-device checkpoints (#28118 approach, rejected).
- Acceptance ~0.53–0.62 at both temps (upstream reports 0.79 at p-min 0.7 on the base model) — the Unc fine-tune / base-head mismatch caps gains.
- Box-noise lesson: temp-0.7 runs swing ±20% under CPU contention from other sessions; temp-0 runs are rock-stable (<1% spread). Medians mandatory; never compare across core sets.
- Earlier no-rollback comparison for reference (
data/bench_v4_65k*.json): v4+MTP lost 20–48% vs v3 at ctx 65536 despite 60–75% acceptance.
3. Decision: keep v3 (no draft) on :12702
At the production sampling preset (temp 0.7) MTP is parity — no rollout on speed alone. At temp 0
(greedy/code/agentic) n=3 gives a stable +24% → rollout candidate only if greedy workloads matter.
Both user presets are acceptance-hostile anyway: thinking runs at temp 1.0 (maximum entropy), and
instruct’s presence 1.5 distorts the target distribution away from draftable argmax paths.
data/DECIDE.json ranking (warmed mean gen t/s): v3_65k 56.9 > v3_200k 55.6 (shipped) >
mtp_v4_65k 39.2 > other MTP variants 35.8–38.5.
4. The 200K validation that never ran — midnight incident post-mortem
out/PLAN_CTX200K.md specified a 4-phase validation (fit ladder 204800→65536 with a 1.5 GiB/GPU
floor, speed A/B under the two user presets, cyber suite, 198K qualification) with trap-guaranteed
production restore, scheduled one-shot via cron 0 0 8 9 * (2026-09-08 00:00 NZST). The cron
fired at 00:00:01 exactly; the failure was environmental, and the runner caught it correctly:
| Time | Event | Evidence |
|---|---|---|
| ~19:45 | Production :12702 stopped externally (SIGTERM, graceful exit) |
v3_sharp.log ends in cleaning up before exit |
| ~19:48 | Third party reorganized WORKSTATION/models: 33 shards consolidated into one 94.5 GB file; deleted the Q4_K_M draft, the download cache, and the UD-IQ4_XS tree (~110 GB freed) |
mtimes; draft ls fails; df 141→251 GB free |
| 20:20 | Another session started a manual proxy on :12434 (11 models, no allowlist) → systemd unit crash-loops on “port already in use” every 10 s |
pid 1726409 vs systemd journal |
| 00:00:01 | Cron fires; preflight correctly aborts: “sharp template or draft missing” (draft deleted) | ctx200k_run.log |
| 00:00:03 | Trap-restore attempts v3 ×2; both die instantly: the script’s MODEL_PATH glob finds 2 candidates (mmproj + consolidated single file) → die; blind 15-min health waits ×2 |
ctx200k_v3restored.log |
| 00:30 | Runner declares PRODUCTION DOWN + Telegram alert (correct given what it could see) | ctx200k_run.log |
Post-mortem findings (runner hardening lessons):
- Glob-based
MODEL_PATHis fragile — a directory with 2 GGUF candidates (weights + mmproj) must resolve to the entrypoint explicitly, else fail fast, not after 30 min of blind waiting. - Health waits should poll with early-exit on dead child processes, not sleep the full window.
- Preflight was the hero: validating artifacts before stopping production is what prevented a 2.5 h wasted downtime window — but the draft-file check ran before checking production was already down; ordering preflight checks by blast radius matters.
- Duplicate log lines (tee + cron redirect) are cosmetic but double every diagnostic.
- Shared-box reality: another actor’s disk cleanup can invalidate a scheduled job’s inputs hours before it fires. Scheduled jobs should re-verify inputs and re-check assumptions at fire time.
Resolution: production was restored externally the next day — :12702 healthy on the consolidated
single-file GGUF + mmproj, with the sharp chat template (--chat-template-file), pid verified
2026-09-09. The Q4_K_M draft is gone (Q8_0 draft survives, 3.9 GB; speed-equivalent per DECIDE.json);
re-export would need the 32 GB selective download again. The cron line remains (fires next on
2027-09-08; remove for hygiene).
5. Router deployment notes (same arc, worth keeping)
- Family-only exposure works via env on the systemd unit:
MODEL_FAMILY_ALLOWLIST+DISABLED_AUXILIARY_MODEL_IDS— but the drop-in must sort last (zz-*.conf) because the stocklaguna-env.confloads/etc/laguna-s21-proxy.envafterwards and overwritesDISABLED_AUXILIARY_MODEL_IDS. - The family JSON uses llama.cpp’s key name
repeat_penalty(notrepetition_penalty). - Sharp template on the backend:
CHAT_TEMPLATE_FILE=WORKSTATION/chat_template.jinjathrough the v3 launch script — required because the baked template ignoresenable_thinkingon this abliterated model (template finding #1 of qwen-next-flash). - Presets deployed (verified live via
/v1/presets): thinking 1.0/0.95/20/min_p 0/pres 0/rep 1.0; instruct 0.7/0.80/20/min_p 0/pres 1.5/rep 1.0.
6. Honest caveats
- All MTP numbers are ctx 32768 single-node medians; production 200K+MTP fit was never measured (prior attempt OOMs, and the validation that would have settled it never ran).
- The draft head v1 attends densely in the MTP block (trunk QSA-indexer pruning not wired into
graph_mtp); numerically a superset, matters only for long-ctx draft cost. - Q4_K_M draft came from q8_0 requant (
--allow-requantizedouble quant); from-BF16 would be marginally better. - Acceptance capped ~0.6 suggests the draft head mismatches the abliterated fine-tune; p-min tuning untested.
- Patches are uncommitted in
WORKSTATION/llama.cpp-flash-nextby instruction; theoutput_hc_normmtp_only fix is worth upstreaming to #27836.
Reproduction
- Export: apply #27836 + §1 fixes, then selective download →
convert_hf_to_gguf.py --mtp --outtype q8_0→llama-quantize --allow-requantize(job log:quant-jobs/mtp-unc-export/logs/). - Smoke:
scripts/smoke_draft_mtp.sh(target + draft, port 12788). - A/B:
scripts/ab_baseline.sh/scripts/ab_draft.sh(taskset-pinned, median-of-5). - Scheduled validation (as designed):
scripts/cron_wrapper_ctx200k.sh→scripts/run_ctx200k_validation.sh; probesscripts/speed_ab_200k.py,scripts/qual_200k.py, suitescripts/bench_ctx200k_mtp.py. - Data:
data/ab_*.json(Track 1),data/decide_*.json+data/DECIDE.json(ranking),data/bench_v3*.json/data/bench_v4*.json(no-rollback comparison).
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.