Plain-language summary of the MTP draft-head experiment. Full data: README.md, data/, scripts/.
What we built
We taught llama.cpp to use Qwen3.8’s built-in “MTP” draft head — a small extra block the model ships with, whose only job is to guess the next few tokens so the big model can check them in one pass. Exporting it took an upstream patch (PR #27836) plus three fixes the patch was missing. It worked: the draft ran with ~61% of its guesses accepted.
The result that matters: it’s a wash at production settings
With the missing rollback fix ported (upstream #28123), draft-MTP went from 2.7x SLOWER than plain decoding to dead even at the production sampling (temp 0.7: 55.3 vs 54.3 t/s). At temperature 0 (greedy, deterministic) it’s a real win: +24% (71 vs 57.4 t/s), rock stable. Deeper drafts (n=6) are always worse.
Why parity and not a win? Two reasons the data points to: acceptance is ~0.6 here vs 0.79 upstream — the abliterated fine-tune drifts from the base model’s draft head — and the winning presets are hostile to speculative decoding anyway: thinking runs at temp 1.0 (maximum randomness), and instruct’s presence penalty pushes sampling off the draft’s most-likely paths. So we kept the plain server. MTP stays in the drawer for greedy workloads.
The operations lesson: midnight taught us more than the benchmark
We scheduled the 200K-context validation to run unattended at midnight. The cron fired on the second. The test still never ran — because earlier that evening, another session (a) stopped the production server, (b) consolidated the 33 model shards into one file and deleted the draft model our preflight checks on, and (c) started its own proxy on our port.
The runner did its job: preflight caught the missing draft BEFORE stopping production and sent the Telegram alert. The restore path then failed for a silly reason — it picked the model file by glob, found two candidates (weights + vision projector), and gave up after two blind 15-minute waits instead of failing fast. Four hardening rules came out of it: explicit paths over globs, poll-with-early-exit instead of fixed sleeps, order preflight checks by blast radius, and assume shared-box cleanup can invalidate your job’s inputs hours before it fires.
Numbers to remember
| claim | number |
|---|---|
| Draft acceptance (smoke, temp 0.7) | 0.615, mean 2.83 tokens |
| MTP + rollback @ temp 0.7 | 55.3 t/s = parity with 54.3 baseline |
| MTP + rollback @ temp 0 | 71.0 t/s = +24% over 57.4 |
| MTP without rollback | 21 t/s (2.7x slower — the fix that mattered) |
| n=6 drafts | always worse than n=3 |
| Server kept | v3, no draft, 55.6 t/s warmed at 200K ctx |
Honest caveats
- Temp-0.7 benchmark numbers on this box swing ±20% under CPU contention; temp-0 numbers are stable. Medians only, same-core comparisons only.
- The 200K + MTP fit question was never answered — the validation that would have settled it is the one that got cancelled by the incident.
- Draft and base differ by exactly 1 token on divergent runs (verify-batch numerics) — harmless, but it means draft runs are not bit-identical to plain runs even at temp 0.