Back to the experimentSupporting notebook

Findings in simple terms

Source: qwen38/flash-next-unc-mtp-draft-2026-09-07/FINDINGS.md · revision 6550ead3945b

Plain-language summary of the MTP draft-head experiment. Full data: README.md, data/, scripts/.

What we built

We taught llama.cpp to use Qwen3.8’s built-in “MTP” draft head — a small extra block the model ships with, whose only job is to guess the next few tokens so the big model can check them in one pass. Exporting it took an upstream patch (PR #27836) plus three fixes the patch was missing. It worked: the draft ran with ~61% of its guesses accepted.

The result that matters: it’s a wash at production settings

With the missing rollback fix ported (upstream #28123), draft-MTP went from 2.7x SLOWER than plain decoding to dead even at the production sampling (temp 0.7: 55.3 vs 54.3 t/s). At temperature 0 (greedy, deterministic) it’s a real win: +24% (71 vs 57.4 t/s), rock stable. Deeper drafts (n=6) are always worse.

Why parity and not a win? Two reasons the data points to: acceptance is ~0.6 here vs 0.79 upstream — the abliterated fine-tune drifts from the base model’s draft head — and the winning presets are hostile to speculative decoding anyway: thinking runs at temp 1.0 (maximum randomness), and instruct’s presence penalty pushes sampling off the draft’s most-likely paths. So we kept the plain server. MTP stays in the drawer for greedy workloads.

The operations lesson: midnight taught us more than the benchmark

We scheduled the 200K-context validation to run unattended at midnight. The cron fired on the second. The test still never ran — because earlier that evening, another session (a) stopped the production server, (b) consolidated the 33 model shards into one file and deleted the draft model our preflight checks on, and (c) started its own proxy on our port.

The runner did its job: preflight caught the missing draft BEFORE stopping production and sent the Telegram alert. The restore path then failed for a silly reason — it picked the model file by glob, found two candidates (weights + vision projector), and gave up after two blind 15-minute waits instead of failing fast. Four hardening rules came out of it: explicit paths over globs, poll-with-early-exit instead of fixed sleeps, order preflight checks by blast radius, and assume shared-box cleanup can invalidate your job’s inputs hours before it fires.

Numbers to remember

claim number
Draft acceptance (smoke, temp 0.7) 0.615, mean 2.83 tokens
MTP + rollback @ temp 0.7 55.3 t/s = parity with 54.3 baseline
MTP + rollback @ temp 0 71.0 t/s = +24% over 57.4
MTP without rollback 21 t/s (2.7x slower — the fix that mattered)
n=6 drafts always worse than n=3
Server kept v3, no draft, 55.6 t/s warmed at 200K ctx

Honest caveats