On this 72 GB workstation, 8-bit KV cache plus the checkpoint’s own MTP head at a draft window of 3 beat no-MTP short prose. A window of 6 lost. Thinking needed a 16k output cap to finish, not 4k.
Executive summary
- Measured (draft-window sweep, thinking off, n=1): native MTP draft=3 mean short-prose 88.16 tok/s versus 64.67 tok/s with MTP off. Relative gain ((88.16-64.67)/64.67 = 36.3%). Draft=6 was 62.35 tok/s, slower than off.
- Measured (winner, instruct vs thinking, max 16384, n=2 except mid): instruct short-prose medians 78.4 / 81.2 tok/s; thinking 110.2 / 133.2 tok/s when chain-of-thought was long. At concurrency 2 through the OpenAI-compatible proxy, instruct aggregate 111.2 tok/s (about 56 tok/s per request) versus 81.4 tok/s solo.
- Recommended for this box:
CACHE_KV_BITS=8,CACHE_SIZE=400128,ENABLE_MTP=1,NUM_DRAFT_TOKENS=3,MAX_BATCH_SIZE=2, Sharp chat template, proxy output cap 40000. Instruct sampler for non-thinking; thinking sampler for thinking; do not raise the draft window to 6. - Principal caveat: the no-MTP comparison is a single-shot short-prose sweep with thinking off. Long-context (30k–75k) traffic was not re-measured with MTP. Thinking length is unstable.
The problem (why)
Qwen3.8-Flash-Next is a hybrid 125B-A6B model with a 4B multi-token prediction (MTP) head. Three 24 GB RTX 3090s are tight: fp16 KV at 400k tokens already fills the cards, so any extra weights have to come from compressing KV, not from “free VRAM.” The operational questions were: can the native MTP component of the official 3.50 HQ EXL3 checkpoint run on this split; which draft window is fastest; which samplers to pin; and how large an output cap thinking actually needs.
A previous llama.cpp experiment on an Uncensored GGUF sibling used a separate MTP draft GGUF. That is a different stack. This run uses Model.from_config(cfg, component="mtp") on the same 3.50 HQ directory.
The system and experiment (what)
- Hardware: 3× NVIDIA GeForce RTX 3090 24 GB, driver 595.58.03, CUDA 13.2.
- Weights: official Qwen3.8-Flash-Next EXL3 3.50 HQ (n-gram table in RAM). MTP tensors live in the last shard of this tree (6203
mtp.*keys), not the Unc EXL3-4bpw tree and notmtp-Qwen38-Unc-Q8_0.gguf. - Runtime: exllamav3 layer-split autosplit,
GPU_SPLIT=24,24,24,CHUNK_SIZE=256,MAX_BATCH_SIZE=2, Sharpchat_template.jinja. - KV: 400128 tokens, 8-bit (
CACHE_KV_BITS=8). Switching fp16→q8 at this length was measured to free enough VRAM for MTP (~1.27 GiB weights plus draft cache) without cutting context. - Front door: variant proxy on port 12434; backend on 12703. Public ids:
qwen38-unc-exl3-instruct/thinking/medium/xhigh. - Variables: MTP off vs draft N ∈ {2,3,4,6}; instruct vs thinking samplers; output caps 256 / 4096 / 16384; concurrency 1 vs 2.
Instruct sampler: temperature 0.7, top_p 0.80, top_k 20, min_p 0, presence_penalty 1.5, repetition_penalty 1.0, thinking off.
Thinking bench sampler: 1.0 / 0.95 / 20 / 0 / presence 0.0 / 1.0, thinking on. Production thinking on the proxy still forces presence_penalty=1.0.
Method (how)
Draft windows required a full reload (N is fixed at generator init). The proxy was SIGSTOP’d during the sweep so it could not respawn an old env. Each cell used the same three prompts (two short prose tasks, one ~3790-token filler). Token/s is end-to-end wall clock including prefill: new_tokens / elapsed. MTP accept is accepted draft tokens / drafted window tokens from exllamav3 draft_stats.
Later instruct/thinking runs reused the winning process (q8 + MTP d3) and only changed request JSON. Caps 256 and 4096 truncated thinking (finish=length). Cap 16384 allowed finish=stop on every scored thinking cell. Concurrency 2 was two parallel POSTs to port 12434; the backend log showed enqueue job active=2/2.
flowchart LR
Client --> Proxy12434
Proxy12434 --> Backend12703
Backend12703 --> Trunk["3.50 HQ trunk + q8 KV 400k"]
Backend12703 --> MTP["native MTP head draft N"]
Trunk --> Measure["wall tok/s + draft_stats"]
MTP --> Measure
Figure 1. Request path and measurement points. Alt: Client hits the proxy, which forwards to the EXL3 server; that process owns both the trunk+KV and the native MTP head; tok/s and accept are logged per job.
Notice: concurrency 2 is two jobs in one generator, not two processes.
This diagram does not prove placement of individual layers on GPU 0/1/2.
Results
Draft window (thinking off, n=1, prose_tps = mean of two short prompts)
Source: data/draft-window-sweep.json.
| tag | MTP | N | prose | prose2 | mean | mid ~3.8k | accept prose/prose2/mid |
|---|---|---|---|---|---|---|---|
| off | no | 0 | 65.98 | 63.35 | 64.67 | 24.65 | — |
| d2 | yes | 2 | 85.18 | 83.57 | 84.38 | 24.89 | 0.54 / 0.57 / 0.56 |
| d3 | yes | 3 | 82.23 | 94.10 | 88.16 | 27.28 | 0.44 / 0.55 / 0.50 |
| d4 | yes | 4 | 76.61 | 84.81 | 80.71 | 27.46 | 0.37 / 0.41 / 0.40 |
| d6 | yes | 6 | 53.28 | 71.42 | 62.35 | 23.76 | 0.23 / 0.28 / 0.30 |
Chart spec (horizontal bars, tok/s, sort by mean descending): d3 88.16, d2 84.38, d4 80.71, off 64.67, d6 62.35. Source field prose_tps.
Figure 2. Short-prose mean tok/s vs draft window. Alt: draft 3 is highest; draft 6 falls below no MTP.
Notice: the ranking is this prompt pair, n=1, thinking off.
It cannot establish 38k-agentic speed or thinking-mode ranking.
Mid-context in this sweep was first-hit-after-load (~25 tok/s). Warm mid instruct later measured 70.3 tok/s (4k cap) and 77.7 tok/s (16k cap) on the same d3 process.
Instruct vs thinking on d3 (n=2, cap 16384)
Source: data/instruct-vs-thinking-16k.json. Aggregation: median of two reps except mid (n=1).
| mode | task | tok/s median | accept mean | new tokens mean | finish |
|---|---|---|---|---|---|
| instruct | prose | 78.4 | 0.43 | 225 | stop |
| instruct | prose2 | 81.2 | 0.43 | 182 | stop |
| instruct | mid ~3790 | 77.7 | 0.45 | 124 | stop |
| thinking | prose | 110.2 | 0.67 | 3192 | stop (565 and 5819 new) |
| thinking | prose2 | 133.2 | 0.88 | 3418 | stop |
| thinking | mid ~3790 | 131.7 | 0.89 | 2826 | stop |
Cap 256 and 4096: thinking often length with almost no answer. Cap 16384: all scored thinking cells stop. Production output cap was then raised to 40000 so a long CoT cannot hit the 16k wall; that 40k cap was not re-benched.
Concurrency 2 through the proxy (d3, max 512)
Source: data/concurrency-2-12434.json.
| 1 request | 2 requests, each | 2-request aggregate | |
|---|---|---|---|
| instruct | 81.37 tok/s | 54.27 / 57.61 | 111.2 tok/s |
| thinking | 85.75 tok/s (length at 512) |
66.39 / 61.43 | 122.8 tok/s |
Throughput ratio instruct: (111.2 / 81.37 = 1.37\times). Per-request speed fell ~30%.
Why the result looks this way
Observed. Q8 KV at 400k plus MTP loaded (VRAM after d3 load: 23414 / 23034 / 20948 MiB). Accept fell as N rose (≈0.55 at N=2, ≈0.25 at N=6). Thinking accept on long CoT was ≈0.87–0.89, higher than instruct ≈0.43. Backend batched two jobs (active=2/2).
Interpretation (not measured as a causal graph). A larger draft window that is mostly rejected wastes a forward of the MTP block. Thinking traces are more predictable, so the same N=3 head pays off. Short instruct answers (e.g. arith, 4 tokens, accept 0.11) get little or negative help. First request after reload pays prefill/kernel warmup; that is why sweep “mid” looked like 25 tok/s and warm mid later looked like ~78 tok/s.
Operational recommendation
For this official 3.50 HQ EXL3 on 3×3090, keep:
- 8-bit KV, 400128-token cache, batch 2, chunk 256, split 24/24/24.
- Native MTP on,
NUM_DRAFT_TOKENS=3(not 6). - Sharp chat template.
- Instruct public id → instruct sampler. Thinking/medium/xhigh → thinking sampler (proxy still uses presence 1.0).
- Proxy
maximum_completion_tokens=40000so thinking can EOS (16k was enough in the bench; 4k was not).
Smoke after a respawn: GET /health on the backend, one instruct 180-word prompt (expect stop near 80 tok/s warm), one thinking prompt with max_tokens>=16384 (expect stop, not length).
Trade-off: MTP adds ~2.6 GiB-class working set (weights + draft cache + GDN history). q8 is what makes it fit beside 400k context. Draft=3 is a speed pick, not a quality study.
Limits and next experiment
- n=1 on the window sweep; n=2 on 16k short cells; no medians across days.
- No MTP-off thinking run at 16k, so the 110–133 tok/s thinking numbers are not a controlled vs-off comparison.
- No 38k/400k MTP rerun (earlier live fp16-no-MTP traffic at 30–75k was ~50 tok/s; different KV and no MTP).
- Thinking bench used presence 0.0; production thinking uses 1.0.
- 40k output cap is configuration only; not timed.
- Next smallest test: same 16k instruct/thinking suite with MTP off on q8 400k, n=3, plus one 32k-prompt warm decode.
How to reproduce or audit
Workstation-only (weights, 3×3090, exllamav3 env):
## replay draft sweep (SIGSTOPs the proxy; reloads the backend per N)
python3 qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/scripts/sweep_draft.py
python3 qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/scripts/run_mode_bench_16k.py
Portable audit: read the JSON files listed below; recompute prose_tps as the mean of the two short-prose tps fields; confirm accept from mtp_accept.
Evidence appendix
- Entry:
qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/README.md data/draft-window-sweep.json,instruct-vs-thinking.json,instruct-vs-thinking-4k.json,instruct-vs-thinking-16k.json,concurrency-2-12434.json,live-backend.json- Scripts:
sweep_draft.py,run_mode_bench.py,run_mode_bench_4k.py,run_mode_bench_16k.py - MTP: multi-token prediction head used as a speculator. Accept: accepted drafted tokens / drafted tokens. q8 KV: 8-bit quantized key/value cache. Sharp template: the shared Qwen chat template used by the GGUF sibling.
What this does not prove: that draft=3 is best on other GPUs, other quants, Unc weights, or 100k-token agent traces; that thinking is “higher quality”; or that 40k output is faster than 16k.