Back to the experimentSupporting notebook

Tuning Qwen3.8-Flash-Next EXL3 3.50 HQ on 3× RTX 3090: q8 KV, native MTP, draft window 3

Source: qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/BLOG.md · revision 6550ead3945b

On this 72 GB workstation, 8-bit KV cache plus the checkpoint’s own MTP head at a draft window of 3 beat no-MTP short prose. A window of 6 lost. Thinking needed a 16k output cap to finish, not 4k.

Executive summary

The problem (why)

Qwen3.8-Flash-Next is a hybrid 125B-A6B model with a 4B multi-token prediction (MTP) head. Three 24 GB RTX 3090s are tight: fp16 KV at 400k tokens already fills the cards, so any extra weights have to come from compressing KV, not from “free VRAM.” The operational questions were: can the native MTP component of the official 3.50 HQ EXL3 checkpoint run on this split; which draft window is fastest; which samplers to pin; and how large an output cap thinking actually needs.

A previous llama.cpp experiment on an Uncensored GGUF sibling used a separate MTP draft GGUF. That is a different stack. This run uses Model.from_config(cfg, component="mtp") on the same 3.50 HQ directory.

The system and experiment (what)

Instruct sampler: temperature 0.7, top_p 0.80, top_k 20, min_p 0, presence_penalty 1.5, repetition_penalty 1.0, thinking off.
Thinking bench sampler: 1.0 / 0.95 / 20 / 0 / presence 0.0 / 1.0, thinking on. Production thinking on the proxy still forces presence_penalty=1.0.

Method (how)

Draft windows required a full reload (N is fixed at generator init). The proxy was SIGSTOP’d during the sweep so it could not respawn an old env. Each cell used the same three prompts (two short prose tasks, one ~3790-token filler). Token/s is end-to-end wall clock including prefill: new_tokens / elapsed. MTP accept is accepted draft tokens / drafted window tokens from exllamav3 draft_stats.

Later instruct/thinking runs reused the winning process (q8 + MTP d3) and only changed request JSON. Caps 256 and 4096 truncated thinking (finish=length). Cap 16384 allowed finish=stop on every scored thinking cell. Concurrency 2 was two parallel POSTs to port 12434; the backend log showed enqueue job active=2/2.

flowchart LR
  Client --> Proxy12434
  Proxy12434 --> Backend12703
  Backend12703 --> Trunk["3.50 HQ trunk + q8 KV 400k"]
  Backend12703 --> MTP["native MTP head draft N"]
  Trunk --> Measure["wall tok/s + draft_stats"]
  MTP --> Measure

Figure 1. Request path and measurement points. Alt: Client hits the proxy, which forwards to the EXL3 server; that process owns both the trunk+KV and the native MTP head; tok/s and accept are logged per job.
Notice: concurrency 2 is two jobs in one generator, not two processes.
This diagram does not prove placement of individual layers on GPU 0/1/2.

Results

Draft window (thinking off, n=1, prose_tps = mean of two short prompts)

Source: data/draft-window-sweep.json.

tag MTP N prose prose2 mean mid ~3.8k accept prose/prose2/mid
off no 0 65.98 63.35 64.67 24.65 —
d2 yes 2 85.18 83.57 84.38 24.89 0.54 / 0.57 / 0.56
d3 yes 3 82.23 94.10 88.16 27.28 0.44 / 0.55 / 0.50
d4 yes 4 76.61 84.81 80.71 27.46 0.37 / 0.41 / 0.40
d6 yes 6 53.28 71.42 62.35 23.76 0.23 / 0.28 / 0.30

Chart spec (horizontal bars, tok/s, sort by mean descending): d3 88.16, d2 84.38, d4 80.71, off 64.67, d6 62.35. Source field prose_tps.
Figure 2. Short-prose mean tok/s vs draft window. Alt: draft 3 is highest; draft 6 falls below no MTP.
Notice: the ranking is this prompt pair, n=1, thinking off.
It cannot establish 38k-agentic speed or thinking-mode ranking.

Mid-context in this sweep was first-hit-after-load (~25 tok/s). Warm mid instruct later measured 70.3 tok/s (4k cap) and 77.7 tok/s (16k cap) on the same d3 process.

Instruct vs thinking on d3 (n=2, cap 16384)

Source: data/instruct-vs-thinking-16k.json. Aggregation: median of two reps except mid (n=1).

mode task tok/s median accept mean new tokens mean finish
instruct prose 78.4 0.43 225 stop
instruct prose2 81.2 0.43 182 stop
instruct mid ~3790 77.7 0.45 124 stop
thinking prose 110.2 0.67 3192 stop (565 and 5819 new)
thinking prose2 133.2 0.88 3418 stop
thinking mid ~3790 131.7 0.89 2826 stop

Cap 256 and 4096: thinking often length with almost no answer. Cap 16384: all scored thinking cells stop. Production output cap was then raised to 40000 so a long CoT cannot hit the 16k wall; that 40k cap was not re-benched.

Concurrency 2 through the proxy (d3, max 512)

Source: data/concurrency-2-12434.json.

1 request 2 requests, each 2-request aggregate
instruct 81.37 tok/s 54.27 / 57.61 111.2 tok/s
thinking 85.75 tok/s (length at 512) 66.39 / 61.43 122.8 tok/s

Throughput ratio instruct: (111.2 / 81.37 = 1.37\times). Per-request speed fell ~30%.

Why the result looks this way

Observed. Q8 KV at 400k plus MTP loaded (VRAM after d3 load: 23414 / 23034 / 20948 MiB). Accept fell as N rose (≈0.55 at N=2, ≈0.25 at N=6). Thinking accept on long CoT was ≈0.87–0.89, higher than instruct ≈0.43. Backend batched two jobs (active=2/2).

Interpretation (not measured as a causal graph). A larger draft window that is mostly rejected wastes a forward of the MTP block. Thinking traces are more predictable, so the same N=3 head pays off. Short instruct answers (e.g. arith, 4 tokens, accept 0.11) get little or negative help. First request after reload pays prefill/kernel warmup; that is why sweep “mid” looked like 25 tok/s and warm mid later looked like ~78 tok/s.

Operational recommendation

For this official 3.50 HQ EXL3 on 3×3090, keep:

  1. 8-bit KV, 400128-token cache, batch 2, chunk 256, split 24/24/24.
  2. Native MTP on, NUM_DRAFT_TOKENS=3 (not 6).
  3. Sharp chat template.
  4. Instruct public id → instruct sampler. Thinking/medium/xhigh → thinking sampler (proxy still uses presence 1.0).
  5. Proxy maximum_completion_tokens=40000 so thinking can EOS (16k was enough in the bench; 4k was not).

Smoke after a respawn: GET /health on the backend, one instruct 180-word prompt (expect stop near 80 tok/s warm), one thinking prompt with max_tokens>=16384 (expect stop, not length).

Trade-off: MTP adds ~2.6 GiB-class working set (weights + draft cache + GDN history). q8 is what makes it fit beside 400k context. Draft=3 is a speed pick, not a quality study.

Limits and next experiment

How to reproduce or audit

Workstation-only (weights, 3×3090, exllamav3 env):

## replay draft sweep (SIGSTOPs the proxy; reloads the backend per N)
python3 qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/scripts/sweep_draft.py
python3 qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/scripts/run_mode_bench_16k.py

Portable audit: read the JSON files listed below; recompute prose_tps as the mean of the two short-prose tps fields; confirm accept from mtp_accept.

Evidence appendix

What this does not prove: that draft=3 is best on other GPUs, other quants, Unc weights, or 100k-token agent traces; that thinking is “higher quality”; or that 40k output is faster than 16k.