Journal

Finding a local model for the inbox

Bonsai-2-27B PQ2_0 with MTP passed 48/50 tasks with thinking on, at a 15.8-second median on an RTX 3060.

Vlad / experimentos.
The boundary

This is one fixed 50-task inbox workload; hardware, runtime and thinking mode all matter.

Measurements

Recorded result

Thinking-on inbox performance across configurations

Same 50 inbox tasks · fixed sampling · different hardware, quantizations and runtimes

Thinking-on inbox performance across configurationsA Swift-1.5 27B IQ3_S (Win): 94 strict pass (%); B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%); C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%); D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%); E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%); F MiMo-9B MLX 4bit (Mac): 52 strict pass (%); F MiMo-9B MLX 4bit (Mac): 64 strict pass (%); G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%); H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%); I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%); J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%); JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%); L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%); M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%); K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%); KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%); NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%). One observed workload. Each label retains the hardware; forced thinking is a distinct mode.A Swift-1.5 27B IQ3_S (Win)94A Swift-1.5 27B IQ3_S (Win): 94 strict pass (%)B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 94 strict pass (%)C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 90 strict pass (%)D Gemma4-12B Q4_K_M+MTP (RTX 3060)76D Gemma4-12B Q4_K_M+MTP (RTX 3060): 76 strict pass (%)E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 32 strict pass (%)F MiMo-9B MLX 4bit (Mac)52F MiMo-9B MLX 4bit (Mac): 52 strict pass (%)F MiMo-9B MLX 4bit (Mac)64F MiMo-9B MLX 4bit (Mac): 64 strict pass (%)G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 10 strict pass (%)H Bonsai-2-27B MLX 2bit (Mac)94H Bonsai-2-27B MLX 2bit (Mac): 94 strict pass (%)I Bonsai-2-27B-Abl PQ2_0 (Mac)90I Bonsai-2-27B-Abl PQ2_0 (Mac): 90 strict pass (%)J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 96 strict pass (%)JMac Bonsai-2-27B PQ2_0+MTP (Mac)96JMac Bonsai-2-27B PQ2_0+MTP (Mac): 96 strict pass (%)L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 92 strict pass (%)M ZDTaichu5.0-9B MLX 8bit (Mac,text-only)82M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 82 strict pass (%)K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPUdraft)96K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 96 strict pass (%)KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 98 strict pass (%)NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP(Mac)86NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 86 strict pass (%)0strict pass (%)
  1. A Swift-1.5 27B IQ3_S (Win)94
  2. B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94
  3. C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90
  4. D Gemma4-12B Q4_K_M+MTP (RTX 3060)76
  5. E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32
  6. F MiMo-9B MLX 4bit (Mac)52
  7. F MiMo-9B MLX 4bit (Mac)64
  8. G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10
  9. H Bonsai-2-27B MLX 2bit (Mac)94
  10. I Bonsai-2-27B-Abl PQ2_0 (Mac)90
  11. J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96
  12. JMac Bonsai-2-27B PQ2_0+MTP (Mac)96
  13. L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92
  14. M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)82
  15. K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)96
  16. KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98
  17. NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)86

strict pass (%)

One observed workload. Each label retains the hardware; forced thinking is a distinct mode.

View data & source
Thinking-on inbox performance across configurations · strict pass (%)
ConfigurationValue
A Swift-1.5 27B IQ3_S (Win)94
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)94
C Qwen3.8-27B GSQ-RCO IQ3_S (Win)90
D Gemma4-12B Q4_K_M+MTP (RTX 3060)76
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)32
F MiMo-9B MLX 4bit (Mac)52
F MiMo-9B MLX 4bit (Mac)64
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)10
H Bonsai-2-27B MLX 2bit (Mac)94
I Bonsai-2-27B-Abl PQ2_0 (Mac)90
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)96
JMac Bonsai-2-27B PQ2_0+MTP (Mac)96
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)92
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)82
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)96
KMac Bonsai-2 PQ2_0+DFlash2 (Mac)98
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)86

Origin: reported aggregate table. Revision 6550ead3945b.

Download measurement metadata (JSON)
Recorded result

The cost of the same inbox workload

Thinking-on and forced-thinking rows · same 50 tasks

The cost of the same inbox workloadA Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / task; B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / task; C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / task; D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / task; E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / task; F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / task; F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / task; G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / task; H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / task; I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / task; J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / task; JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / task; L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / task; M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / task; K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / task; KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / task; NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task. Median client wall time. A higher pass rate does not automatically give a faster workflow.A Swift-1.5 27B IQ3_S (Win)51.2A Swift-1.5 27B IQ3_S (Win): 51.2 median seconds / taskB Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6B Bonsai-2-27B-Abl PQ2_0 (RTX 3060): 26.6 median seconds / taskC Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9C Qwen3.8-27B GSQ-RCO IQ3_S (Win): 59.9 median seconds / taskD Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0D Gemma4-12B Q4_K_M+MTP (RTX 3060): 8.0 median seconds / taskE Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0E Xing4.0-29B-A4B IQ3_XXS (RTX 3060): 69.0 median seconds / taskF MiMo-9B MLX 4bit (Mac)6.7F MiMo-9B MLX 4bit (Mac): 6.7 median seconds / taskF MiMo-9B MLX 4bit (Mac)17.4F MiMo-9B MLX 4bit (Mac): 17.4 median seconds / taskG Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060): 94.4 median seconds / taskH Bonsai-2-27B MLX 2bit (Mac)87.0H Bonsai-2-27B MLX 2bit (Mac): 87.0 median seconds / taskI Bonsai-2-27B-Abl PQ2_0 (Mac)81.7I Bonsai-2-27B-Abl PQ2_0 (Mac): 81.7 median seconds / taskJ54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060): 15.8 median seconds / taskJMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1JMac Bonsai-2-27B PQ2_0+MTP (Mac): 67.1 median seconds / taskL ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9L ThinkingCap-Qwen3.8-27B IQ3_S (Mac): 81.9 median seconds / taskM ZDTaichu5.0-9B MLX 8bit (Mac,text-only)65.0M ZDTaichu5.0-9B MLX 8bit (Mac, text-only): 65.0 median seconds / taskK54 Bonsai-2 PQ2_0+DFlash2 (3060, CPUdraft)52.0K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft): 52.0 median seconds / taskKMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6KMac Bonsai-2 PQ2_0+DFlash2 (Mac): 71.6 median seconds / taskNMac HauhauCS Qwen3.8-27B IQ3_XS+MTP(Mac)135.5NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac): 135.5 median seconds / task0median seconds / task
  1. A Swift-1.5 27B IQ3_S (Win)51.2
  2. B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6
  3. C Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9
  4. D Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0
  5. E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0
  6. F MiMo-9B MLX 4bit (Mac)6.7
  7. F MiMo-9B MLX 4bit (Mac)17.4
  8. G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4
  9. H Bonsai-2-27B MLX 2bit (Mac)87.0
  10. I Bonsai-2-27B-Abl PQ2_0 (Mac)81.7
  11. J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8
  12. JMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1
  13. L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9
  14. M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)65.0
  15. K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)52.0
  16. KMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6
  17. NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)135.5

median seconds / task

Median client wall time. A higher pass rate does not automatically give a faster workflow.

View data & source
The cost of the same inbox workload · median seconds / task
ConfigurationValue
A Swift-1.5 27B IQ3_S (Win)51.2
B Bonsai-2-27B-Abl PQ2_0 (RTX 3060)26.6
C Qwen3.8-27B GSQ-RCO IQ3_S (Win)59.9
D Gemma4-12B Q4_K_M+MTP (RTX 3060)8.0
E Xing4.0-29B-A4B IQ3_XXS (RTX 3060)69.0
F MiMo-9B MLX 4bit (Mac)6.7
F MiMo-9B MLX 4bit (Mac)17.4
G Hermes-4-14B-OBL EXL3 4.0 (RTX 3060)94.4
H Bonsai-2-27B MLX 2bit (Mac)87.0
I Bonsai-2-27B-Abl PQ2_0 (Mac)81.7
J54 Bonsai-2-27B PQ2_0+MTP (RTX 3060)15.8
JMac Bonsai-2-27B PQ2_0+MTP (Mac)67.1
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac)81.9
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only)65.0
K54 Bonsai-2 PQ2_0+DFlash2 (3060, CPU draft)52.0
KMac Bonsai-2 PQ2_0+DFlash2 (Mac)71.6
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac)135.5

Origin: reported aggregate table. Revision 6550ead3945b.

Download measurement metadata (JSON)

The question

Inbox work needs accurate structured answers and reasonable response times. The benchmark ran the same 50 tasks across local model, runtime and hardware configurations, with thinking on and off, to make that tradeoff visible.

From the original notebook

Which local model should do email triage, extraction and drafting on the home hardware? 17 model/runtime configurations were run through the same 50 tasks (easy 10 + hard 20 + ultra 20), thinking on and off.

Winner: Bonsai-2-27B PQ2_0 + grafted MTP head (J54) on the .54 RTX 3060: 96% (48/50) strict pass with thinking on, median 15.8 s per task, max 43.3 s (nothing over 120 s), about 52 tok/s decode. It is the only >=94% configuration that is also fast.

Hardware

Name Hardware Runtime(s) used
.54 Linux (Ubuntu 24.04), RTX 3060 12 GB, 8 cores, 31 GB RAM. Also runs a vLLM embedder (~2.5 GB VRAM) at all times PrismML llama.cpp (+MTP patch), stock llama.cpp, Xing fork, TabbyAPI/EXL3
Windows RTX 3060 12 GB (no other GPU load) llama.cpp b11177 CUDA 12.4 / PrismML Windows build
Mac Apple M5, 24 GB unified PrismML llama.cpp Metal, stock llama.cpp, mlx-lm, mlx-vlm

Hosts are referred to as .54 / Windows / Mac. code/bench.py reads WIN_HOST, LINUX54_HOST and MAC_HOST from the environment; no real addresses are included.

Final table (all 50 tasks, strict grading)

Model Mode Pass % (strict) easy hard ultra Median s Max s >120 s Decode tok/s Trunc VRAM / memory
A Swift-1.5 27B IQ3_S (Win) off 66% (33/50) 9/10 16/20 8/20 8.9 26.4 0 14.4 0 12070 MiB VRAM (218 free)
A Swift-1.5 27B IQ3_S (Win) on 94% (47/50) 10/10 19/20 18/20 51.2 210.1 6 14.6 1 12070 MiB VRAM (218 free)
B Bonsai-2-27B-Abl PQ2_0 (.54) off 72% (36/50) 10/10 14/20 12/20 4.8 15.5 0 33.6 0 10135 MiB VRAM (2153 free)
B Bonsai-2-27B-Abl PQ2_0 (.54) on 94% (47/50) 10/10 17/20 20/20 26.6 68.2 0 33.6 0 10135 MiB VRAM (2153 free)
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) off 68% (34/50) 10/10 16/20 8/20 9.4 55.6 0 14.8 2 12070 MiB VRAM (218 free)
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) on 90% (45/50) 10/10 19/20 16/20 59.9 207.4 9 14.9 5 12070 MiB VRAM (218 free)
D Gemma4-12B Q4_K_M+MTP (.54) off 66% (33/50) 9/10 16/20 8/20 2.4 6.3 0 62.0 0 11187 MiB VRAM (1101 free)
D Gemma4-12B Q4_K_M+MTP (.54) on 76% (38/50) 8/10 16/20 14/20 8.0 40.6 0 51.6 1 11187 MiB VRAM (1101 free)
E Xing4.0-29B-A4B IQ3_XXS (.54) off 36% (18/50) 9/10 7/20 2/20 7.3 39.3 0 21.1 1 10921 MiB VRAM (1367 free)
E Xing4.0-29B-A4B IQ3_XXS (.54) on 32% (16/50) 9/10 6/20 1/20 69.0 157.1 17 20.7 13 10921 MiB VRAM (1367 free)
F MiMo-9B MLX 4bit (Mac) off 52% (26/50) 10/10 14/20 2/20 6.6 18.8 0 18.7 0 ~7.2-9.4 GB unified (Mac)
F MiMo-9B MLX 4bit (Mac) on 52% (26/50) 9/10 14/20 3/20 6.7 176.2 3 20.6 3 ~7.2-9.4 GB unified (Mac)
F MiMo-9B MLX 4bit (Mac) on_forced 64% (32/50) 9/10 16/20 7/20 17.4 162.9 2 18.9 2 ~7.2-9.4 GB unified (Mac)
G Hermes-4-14B-OBL EXL3 4.0 (.54) off 8% (4/50) 8/10 13/20 4/20 25.5 32.1 0 32.8 46 10669 MiB VRAM (1619 free; Q6 cache, chunk 512)
G Hermes-4-14B-OBL EXL3 4.0 (.54) on 10% (5/50) 8/10 10/20 5/20 94.4 106.5 0 13.1 40 10669 MiB VRAM (1619 free; Q6 cache, chunk 512)
H Bonsai-2-27B MLX 2bit (Mac) off 62% (31/50) 10/10 13/20 8/20 14.9 50.4 0 9.9 0 9966 MB footprint (Mac)
H Bonsai-2-27B MLX 2bit (Mac) on 94% (47/50) 10/10 18/20 19/20 87.0 260.2 16 9.8 0 9966 MB footprint (Mac)
I Bonsai-2-27B-Abl PQ2_0 (Mac) off 72% (36/50) 10/10 14/20 12/20 12.0 51.3 0 11.6 0 ~10.0 GB footprint (Mac)
I Bonsai-2-27B-Abl PQ2_0 (Mac) on 90% (45/50) 10/10 16/20 19/20 81.7 246.7 17 11.0 0 ~10.0 GB footprint (Mac)
J54 Bonsai-2-27B PQ2_0+MTP (.54) off 66% (33/50) 10/10 14/20 9/20 3.8 9.7 0 53.0 0 10963 MiB VRAM (1325 free)
J54 Bonsai-2-27B PQ2_0+MTP (.54) on 96% (48/50) 10/10 18/20 20/20 15.8 43.3 0 52.0 0 10963 MiB VRAM (1325 free)
JMac Bonsai-2-27B PQ2_0+MTP (Mac) off 66% (33/50) 10/10 14/20 9/20 13.1 50.8 0 12.1 0 ~10.0 GB footprint (Mac)
JMac Bonsai-2-27B PQ2_0+MTP (Mac) on 96% (48/50) 10/10 19/20 19/20 67.1 172.6 10 11.3 0 ~10.0 GB footprint (Mac)
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) off 68% (34/50) 10/10 16/20 8/20 23.1 134.2 2 6.4 2 12.6 GB weights, footprint ~9.9 GB + paging (Mac)
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) on 92% (46/50) 10/10 18/20 18/20 81.9 310.9 14 6.2 0 12.6 GB weights, footprint ~9.9 GB + paging (Mac)
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) off 46% (23/50) 10/10 11/20 2/20 6.7 33.2 0 13.5 0 12 GB footprint peak (Mac)
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) on 82% (41/50) 10/10 17/20 14/20 65.0 227.1 10 13.3 1 12 GB footprint peak (Mac)
K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) off 66% (33/50) 10/10 14/20 9/20 8.5 31.5 0 14.8 0 11203 MiB VRAM (1085 free; draft on CPU)
K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) on 96% (48/50) 10/10 19/20 19/20 52.0 243.7 7 14.1 1 11203 MiB VRAM (1085 free; draft on CPU)
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) off 66% (33/50) 10/10 14/20 9/20 13.1 57.4 0 12.0 0 ~9.3 GB weights, footprint ~8.5-10 GB (Mac)
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) on 98% (49/50) 10/10 19/20 20/20 71.6 172.2 11 11.6 0 ~9.3 GB weights, footprint ~8.5-10 GB (Mac)
N54 HauhauCS Qwen3.8-27B IQ3_XS+MTP (.54, partial offload) off 60% (30/50) 10/10 12/20 8/20 48.4 184.8 8 2.0 0 11221 MiB VRAM (1067 free; -ngl 44/65, rest on CPU)
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) off 62% (31/50) 10/10 13/20 8/20 17.6 60.7 0 7.6 0 Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap)
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) on 86% (43/50) 10/10 18/20 15/20 135.5 467.6 27 6.8 5 Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap)

Notes: sampling for every request was temperature 0.2, top_p 0.95, top_k 20, seed 42. Thinking ON used max_tokens 3000; OFF used 800. Requests ran one at a time. Wall time is measured at the client. Full per-task detail, repeat runs and caveats are in summary.md.

Winning configuration: exact J launch command (.54)

Binary: PrismML llama.cpp (prism-b10685 source) plus scripts/0001-qwen35-mtp-hadamard-inverse.patch, built with scripts/build_prismml_mtp.sh (CUDA 12.8, sm_86). Model: Bonsai-2-27B-PQ2_0-MTP.gguf from decent-jawfish/bonsai-2-27b-mtp (sha256 78df4279d40ebebdccfd2dae0e9d4847afee52e94f48f3542ae9437220dbd847).

As benchmarked (2026-09-27, -c 16384):

./llama-server -m Bonsai-2-27B-PQ2_0-MTP.gguf --alias bonsai-27b-pq2-mtp --host 0.0.0.0 --port 8801 \
  -ngl 999 -fa on -c 16384 -np 1 -ctk q8_0 -ctv q8_0 --jinja \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --chat-template-kwargs '{"reasoning_effort": "medium"}' \
  --spec-type draft-mtp --spec-draft-n-max 2

Production since 2026-09-28: same command with -c 28672 ... -ctkd q8_0 -ctvd q8_0 -fit off. That is the largest context that stays fully on the GPU with at least 300 MiB free (see below).

Max-context search on .54 (2026-09-28)

Conditions: fully on the GPU (-ngl 999, -fit off, no CPU offload), draft MTP on (--spec-draft-n-max 2), q8_0 KV for the target and the draft (-ctk/-ctv/-ctkd/-ctvd q8_0), -fa on, -np 1, with the vLLM embedder resident (2548 MiB). A setting only counts if the server loads and a synthetic prompt filling about 90% of the context completes, with at least 300 MiB of GPU memory free at peak. The script is scripts/ctx_test.py and the raw results are in results/ctx_search_54.jsonl. Each extra 2048 tokens of context costs about 46 MiB.

Model -c Loads 90% prompt (tokens) Prompt tok/s Decode tok/s @ 90% fill Decode tok/s (short, JSON) Draft accept VRAM used, loaded (MiB, whole GPU) Min free during prompt (MiB) Counts (>=300 MiB free)
Abliterated-v2 PQ2_0 MTP 16384 OK 14769 (90.1%) 496.8 35.9 56.2 0.92 11059 n/a (first run) no, margin
Bonsai-2-27B PQ2_0 MTP 16384 OK 14769 (90.1%) 492.6 37.1 57.1 0.933 10963 933 yes
Bonsai-2-27B PQ2_0 MTP 24576 OK 22107 (90.0%) 474.8 34.3 56.8 0.933 11333 563 yes
Bonsai-2-27B PQ2_0 MTP 28672 OK 25817 (90.0%) 465.8 33.8 56.7 0.933 11517 379 yes
Bonsai-2-27B PQ2_0 MTP 30720 OK 27636 (90.0%) 460.4 33.4 56.8 0.933 11609 287 no, margin
Bonsai-2-27B PQ2_0 MTP 32768 OK 29484 (90.0%) 456.3 32.5 56.9 0.933 11701 195 no, margin
Abliterated-v2 PQ2_0 MTP 26624 OK 23977 (90.1%) 470.4 33.6 56.0 0.92 11521 375 yes
Abliterated-v2 PQ2_0 MTP 28672 OK 25817 (90.0%) 465.0 33.6 55.7 0.92 11613 283 no, margin
  • Bonsai-2-27B PQ2_0 MTP: max safe -c 28672 (379 MiB free at peak). 30720 and 32768 also ran but left only 287 and 195 MiB free.
  • Abliterated-v2 PQ2_0 MTP: max safe -c 26624 (375 MiB free). The file is about 100 MB larger, so it fits one step less.
  • For comparison, the Windows 3060 has no embedder and reaches -c 77824.

Abliterated variant with MTP

The bench’s “B” model was Ternary-Bonsai-2-27B-Abliterated-PQ2_0.gguf (no MTP). Kept on .54 for later use (not serving) is BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF, file Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf (sha256 a4e4c7b578131595c1694354bd6c74d00920df1f5082647a6126899753ebebf8, verified). Its model card says v2 fixes v1’s thinking-mode blank answers. It loads with --spec-type draft-mtp on the same binary: 56 tok/s on a short JSON generation (draft acceptance 0.92) and 33.6 tok/s at 90% of a 26.6K context. It was not run through the 50-task suite.

Layout

  • code/: bench.py (runner), tasks.py / tasks_hard.py / tasks_ultra.py (prompts + graders), summarize.py, combined.py, final_table.py, vram_stats.py, speedtest.py
  • results/: raw JSONL for every run (results_<X>.jsonl = easy, results_hard_<X>.jsonl, results_ultra_<X>.jsonl; letter = model key in bench.py), plus ctx_search_54.jsonl
  • summary.md: full generated report (per-suite and per-task tables, repeat runs, caveats); final_table.md: the table above
  • scripts/: MTP patch, build script, and the context-search harness

To reproduce: cd code && LINUX54_HOST=<host> python3 bench.py --model J --modes off,on --suite tasks_hard --out results_hard_J.jsonl (--suite is tasks, tasks_hard or tasks_ultra; --base overrides the endpoint). summarize.py and combined.py read the results*.jsonl files from the working directory.

Clean-up after the bench (2026-09-28)

Only Bonsai models were kept (base MTP on .54, Windows and Mac, plus the abliterated v2 MTP on .54). The Swift/GSQ/Gemma/Xing/MiMo/Hermes/ThinkingCap/ZDTaichu/HauhauCS weights were deleted.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS