Host: research workstation — 3× RTX 3090 (24 GB), driver 595.58.03, CUDA 13.2
Model: WORKSTATION/Qwen3.8-Flash-Next-exl3-3.50bpw_hq_h6_ng6 served as qwen38-flash-next-exl3-350bpw on :12703
Draft: native component="mtp" from this same checkpoint (6203 mtp.* tensors in shard 8). Not Qwen3.8-Flash-Next-Unc-exl3-4bpw and not mtp-Qwen38-Unc-Q8_0.gguf.
Runtime: exllamav3 layer-split autosplit, CACHE_SIZE=400128, CACHE_KV_BITS=8, GPU_SPLIT=24,24,24, MAX_BATCH_SIZE=2, CHUNK_SIZE=256, NGRAM_RAM=1, ENABLE_MTP=1.
Proxy: qwen36_variant_proxy.py on :12434, family qwen38-flash-next-exl3-350bpw.
Two measurements:
- Draft window sweep at instruct-like sampling (
temp=0.7 top_p=0.8, thinking off). - Same prompts on the winning window (
NUM_DRAFT_TOKENS=3) with instruct vs thinking production-style samplers.
Raw JSON: data/draft-window-sweep.json, data/instruct-vs-thinking.json (cap 256), data/instruct-vs-thinking-4k.json (cap 4096/2048), data/instruct-vs-thinking-16k.json (cap 16384).
Replay: scripts/sweep_draft.py, scripts/run_mode_bench.py, scripts/run_mode_bench_4k.py, scripts/run_mode_bench_16k.py.
1. Draft window (q8 KV, thinking off)
Prompts: prose (KV-quant explanation, max 256), prose2 (attention decode, max 256), mid (~3790-token filler + summary, max 192). Warmup discarded. One shot per cell. prose_tps = mean of the two short prose runs.
| tag | MTP | draft N | prose tok/s | prose2 tok/s | mid tok/s | accept prose / prose2 / mid | VRAM used MiB (0/1/2) |
|---|---|---|---|---|---|---|---|
| off | no | 0 | 65.98 | 63.35 | 24.65 | — | 23030 / 22372 / 19826 |
| d2 | yes | 2 | 85.18 | 83.57 | 24.89 | 0.54 / 0.57 / 0.56 | 23354 / 22952 / 20888 |
| d3 | yes | 3 | 82.23 | 94.10 | 27.28 | 0.44 / 0.55 / 0.50 | 23414 / 23034 / 20948 |
| d4 | yes | 4 | 76.61 | 84.81 | 27.46 | 0.37 / 0.41 / 0.40 | 23474 / 23134 / 21010 |
| d6 | yes | 6 | 53.28 | 71.42 | 23.76 | 0.23 / 0.28 / 0.30 | 23616 / 23276 / 21130 |
Winner for short prose mean: d3 at 88.16 tok/s (d2 84.38, d4 80.71, off 64.67, d6 62.35).
Live backend left on q8 + MTP draft=3. Mid-context (~3.8k) stays ~25–27 tok/s regardless of N; MTP does not buy much there.
2. Instruct vs thinking on the live winner (draft=3, q8, 400128)
Samplers as requested (not the proxy thinking presence_penalty=1.0 preset):
| mode | temperature | top_p | top_k | min_p | presence_penalty | repetition_penalty | enable_thinking |
|---|---|---|---|---|---|---|---|
| instruct | 0.7 | 0.80 | 20 | 0.0 | 1.5 | 1.0 | false |
| thinking | 1.0 | 0.95 | 20 | 0.0 | 0.0 | 1.0 | true |
| mode | prose tok/s (accept) | prose2 tok/s (accept) | mid tok/s (accept) | prose mean |
|---|---|---|---|---|
| instruct | 84.64 (0.46) | 76.15 (0.38) | 23.55 (0.36) | 80.40 |
| thinking | 100.28 (0.57) | 90.46 (0.50) | 33.99 (0.47) | 95.37 |
Thinking prose/prose2 both hit max_new_tokens=256 with ~1250 chars of reasoning_content and little/no visible answer — tok/s here is mostly draft-friendly chain-of-thought, not finished instruct answers. Instruct produced full answers and stopped on EOS.
3. Better run — cap 4096 (mid 2048), n=2, wait for EOS
Same samplers as §2, same live backend (q8 + MTP d3). Short prompts and arith: max_tokens=4096. Mid: 2048. Two reps except mid/warmup. Timed end-to-end including prefill; kernels already warm.
| mode | task | tps median | accept mean | new tokens mean | finish |
|---|---|---|---|---|---|
| instruct | arith 17×19 | 13.6 | 0.11 | 4 | stop (answer 323) |
| instruct | prose | 79.4 | 0.40 | 221 | stop |
| instruct | prose2 | 78.2 | 0.42 | 171 | stop |
| instruct | mid ~3790 | 70.3 | 0.38 | 131 | stop |
| thinking | arith 17×19 | 93.4 | 0.70 | 124 | stop (answer present) |
| thinking | prose | 132.1 | 0.87 | 4096 | length both reps (CoT ~11k chars, short answer) |
| thinking | prose2 | 133.2 | 0.88 | 3054 | stop |
| thinking | mid ~3790 | 129.6 | 0.88 | 2048 | length (CoT only, 0 content chars) |
Instruct short-prose median ~78.8 tok/s, all EOS. Mid after warmup is 70 tok/s, not the cold ~24 tok/s from §1–2.
Thinking decode is ~130 tok/s at accept ~0.87, but a 180-word ask still does not finish inside 4096 new tokens (rambling CoT). prose2 did EOS around 2.5–3.6k. Arith finishes and is faster than instruct because instruct emits 4 tokens (poor MTP) while thinking drafts a long trace.
4. Cap 16384 — thinking can actually EOS
Same samplers, same backend, max_tokens=16384 on every scored prompt, n=2 except mid.
| mode | task | tps median | accept mean | new tokens mean | finish |
|---|---|---|---|---|---|
| instruct | arith | 13.6 | 0.11 | 4 | stop |
| instruct | prose | 78.4 | 0.43 | 225 | stop |
| instruct | prose2 | 81.2 | 0.43 | 182 | stop |
| instruct | mid ~3790 | 77.7 | 0.45 | 124 | stop |
| thinking | arith | 96.4 | 0.79 | 74 | stop |
| thinking | prose | 110.2 | 0.67 | 3192 | stop (rep0 5819 tok / 133 t/s; rep1 565 tok / 87 t/s) |
| thinking | prose2 | 133.2 | 0.88 | 3418 | stop |
| thinking | mid ~3790 | 131.7 | 0.89 | 2826 | stop (content 998 chars) |
At 16k every scored thinking run hit stop, not length. Instruct stays ~78–81 tok/s short / 77.7 tok/s mid. Thinking decode stays ~110–133 tok/s when CoT is long; the short thinking prose rep drops to 87 tok/s with accept 0.44.
Interpretation (not a universal claim)
- Native MTP of this 3.50 HQ checkpoint is real and faster than no-MTP on short prose at d2–d4.
- d3 is the keep setting for this box: best short-prose mean, mid tied with d4, accept ~0.5.
- d6 is a net loss vs no MTP.
- Thinking looks faster because reasoning tokens accept better (~0.8–0.9 vs ~0.43). At 16k that speed is on finished answers, not only truncated CoT.
- Warm mid-ctx instruct is ~78 tok/s in the 16k run; the ~24 tok/s figures in §1–2 were first-hit-after-load.
- Thinking CoT length is unstable (565 vs 5819 new tokens on the same prose prompt).
- Earlier live traffic at 30–75k without MTP was ~50 tok/s; that long-ctx point was not re-measured in this file.
Limits
- Draft-window sweep is n=1 per cell. 4k/16k runs are n=2 on short/arith, n=1 on mid.
- Short and ~4k prompts only; no 38k/400k rerun.
- 256 and 4096 caps truncated thinking; 16384 did not.
- Thinking
presence_penalty=0.0as specified for this bench; production thinking preset on:12434still forcespresence_penalty=1.0. - Production
maximum_completion_tokens/DEFAULT_MAX_TOKENSraised to 40000 after the 16k run (not re-timed). Sharpchat_template.jinjaconfirmed on the live process.
5. Concurrency 2 on :12434 (d3, max 512)
Source: data/concurrency-2-12434.json. Backend logged active=2/2.
| 1 req tok/s | each of 2 | aggregate | |
|---|---|---|---|
| instruct | 81.37 | 54.27 / 57.61 | 111.2 |
| thinking | 85.75 (length) |
66.39 / 61.43 | 122.8 |