Back to the experimentSupporting notebook

Official Flash-Next EXL3 3.50 HQ: q8 KV + native MTP draft window (2026-09-16)

Source: qwen-next-flash/exl3-350hq-mtp-draft-q8-2026-09-16/README.md · revision 6550ead3945b

Host: research workstation — 3× RTX 3090 (24 GB), driver 595.58.03, CUDA 13.2
Model: WORKSTATION/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 served as qwen38-flash-next-exl3-350bpw on :12703
Draft: native component="mtp" from this same checkpoint (6203 mtp.* tensors in shard 8). Not Qwen3.8-Flash-Next-Unc-exl3-4bpw and not mtp-Qwen38-Unc-Q8_0.gguf.
Runtime: exllamav3 layer-split autosplit, CACHE_SIZE=400128, CACHE_KV_BITS=8, GPU_SPLIT=24,24,24, MAX_BATCH_SIZE=2, CHUNK_SIZE=256, NGRAM_RAM=1, ENABLE_MTP=1.
Proxy: qwen36_variant_proxy.py on :12434, family qwen38-flash-next-exl3-350bpw.

Two measurements:

  1. Draft window sweep at instruct-like sampling (temp=0.7 top_p=0.8, thinking off).
  2. Same prompts on the winning window (NUM_DRAFT_TOKENS=3) with instruct vs thinking production-style samplers.

Raw JSON: data/draft-window-sweep.json, data/instruct-vs-thinking.json (cap 256), data/instruct-vs-thinking-4k.json (cap 4096/2048), data/instruct-vs-thinking-16k.json (cap 16384).
Replay: scripts/sweep_draft.py, scripts/run_mode_bench.py, scripts/run_mode_bench_4k.py, scripts/run_mode_bench_16k.py.

1. Draft window (q8 KV, thinking off)

Prompts: prose (KV-quant explanation, max 256), prose2 (attention decode, max 256), mid (~3790-token filler + summary, max 192). Warmup discarded. One shot per cell. prose_tps = mean of the two short prose runs.

tag MTP draft N prose tok/s prose2 tok/s mid tok/s accept prose / prose2 / mid VRAM used MiB (0/1/2)
off no 0 65.98 63.35 24.65 — 23030 / 22372 / 19826
d2 yes 2 85.18 83.57 24.89 0.54 / 0.57 / 0.56 23354 / 22952 / 20888
d3 yes 3 82.23 94.10 27.28 0.44 / 0.55 / 0.50 23414 / 23034 / 20948
d4 yes 4 76.61 84.81 27.46 0.37 / 0.41 / 0.40 23474 / 23134 / 21010
d6 yes 6 53.28 71.42 23.76 0.23 / 0.28 / 0.30 23616 / 23276 / 21130

Winner for short prose mean: d3 at 88.16 tok/s (d2 84.38, d4 80.71, off 64.67, d6 62.35).
Live backend left on q8 + MTP draft=3. Mid-context (~3.8k) stays ~25–27 tok/s regardless of N; MTP does not buy much there.

2. Instruct vs thinking on the live winner (draft=3, q8, 400128)

Samplers as requested (not the proxy thinking presence_penalty=1.0 preset):

mode temperature top_p top_k min_p presence_penalty repetition_penalty enable_thinking
instruct 0.7 0.80 20 0.0 1.5 1.0 false
thinking 1.0 0.95 20 0.0 0.0 1.0 true
mode prose tok/s (accept) prose2 tok/s (accept) mid tok/s (accept) prose mean
instruct 84.64 (0.46) 76.15 (0.38) 23.55 (0.36) 80.40
thinking 100.28 (0.57) 90.46 (0.50) 33.99 (0.47) 95.37

Thinking prose/prose2 both hit max_new_tokens=256 with ~1250 chars of reasoning_content and little/no visible answer — tok/s here is mostly draft-friendly chain-of-thought, not finished instruct answers. Instruct produced full answers and stopped on EOS.

3. Better run — cap 4096 (mid 2048), n=2, wait for EOS

Same samplers as §2, same live backend (q8 + MTP d3). Short prompts and arith: max_tokens=4096. Mid: 2048. Two reps except mid/warmup. Timed end-to-end including prefill; kernels already warm.

mode task tps median accept mean new tokens mean finish
instruct arith 17×19 13.6 0.11 4 stop (answer 323)
instruct prose 79.4 0.40 221 stop
instruct prose2 78.2 0.42 171 stop
instruct mid ~3790 70.3 0.38 131 stop
thinking arith 17×19 93.4 0.70 124 stop (answer present)
thinking prose 132.1 0.87 4096 length both reps (CoT ~11k chars, short answer)
thinking prose2 133.2 0.88 3054 stop
thinking mid ~3790 129.6 0.88 2048 length (CoT only, 0 content chars)

Instruct short-prose median ~78.8 tok/s, all EOS. Mid after warmup is 70 tok/s, not the cold ~24 tok/s from §1–2.
Thinking decode is ~130 tok/s at accept ~0.87, but a 180-word ask still does not finish inside 4096 new tokens (rambling CoT). prose2 did EOS around 2.5–3.6k. Arith finishes and is faster than instruct because instruct emits 4 tokens (poor MTP) while thinking drafts a long trace.

4. Cap 16384 — thinking can actually EOS

Same samplers, same backend, max_tokens=16384 on every scored prompt, n=2 except mid.

mode task tps median accept mean new tokens mean finish
instruct arith 13.6 0.11 4 stop
instruct prose 78.4 0.43 225 stop
instruct prose2 81.2 0.43 182 stop
instruct mid ~3790 77.7 0.45 124 stop
thinking arith 96.4 0.79 74 stop
thinking prose 110.2 0.67 3192 stop (rep0 5819 tok / 133 t/s; rep1 565 tok / 87 t/s)
thinking prose2 133.2 0.88 3418 stop
thinking mid ~3790 131.7 0.89 2826 stop (content 998 chars)

At 16k every scored thinking run hit stop, not length. Instruct stays ~78–81 tok/s short / 77.7 tok/s mid. Thinking decode stays ~110–133 tok/s when CoT is long; the short thinking prose rep drops to 87 tok/s with accept 0.44.

Interpretation (not a universal claim)

Limits

5. Concurrency 2 on :12434 (d3, max 512)

Source: data/concurrency-2-12434.json. Backend logged active=2/2.

1 req tok/s each of 2 aggregate
instruct 81.37 54.27 / 57.61 111.2
thinking 85.75 (length) 66.39 / 61.43 122.8