Which local model should do email triage, extraction and drafting on the home hardware? 17 model/runtime configurations were run through the same 50 tasks (easy 10 + hard 20 + ultra 20), thinking on and off.
Winner: Bonsai-2-27B PQ2_0 + grafted MTP head (J54) on the .54 RTX 3060: 96% (48/50) strict pass with thinking on,
median 15.8 s per task, max 43.3 s (nothing over 120 s), about 52 tok/s decode. It is the only >=94% configuration that is also fast.
Hardware
| Name | Hardware | Runtime(s) used |
|---|---|---|
| .54 | Linux (Ubuntu 24.04), RTX 3060 12 GB, 8 cores, 31 GB RAM. Also runs a vLLM embedder (~2.5 GB VRAM) at all times | PrismML llama.cpp (+MTP patch), stock llama.cpp, Xing fork, TabbyAPI/EXL3 |
| Windows | RTX 3060 12 GB (no other GPU load) | llama.cpp b11177 CUDA 12.4 / PrismML Windows build |
| Mac | Apple M5, 24 GB unified | PrismML llama.cpp Metal, stock llama.cpp, mlx-lm, mlx-vlm |
Hosts are referred to as .54 / Windows / Mac. code/bench.py reads WIN_HOST, LINUX54_HOST and MAC_HOST from the environment; no real addresses are included.
Final table (all 50 tasks, strict grading)
| Model | Mode | Pass % (strict) | easy | hard | ultra | Median s | Max s | >120 s | Decode tok/s | Trunc | VRAM / memory |
|---|---|---|---|---|---|---|---|---|---|---|---|
| A Swift-1.5 27B IQ3_S (Win) | off | 66% (33/50) | 9/10 | 16/20 | 8/20 | 8.9 | 26.4 | 0 | 14.4 | 0 | 12070 MiB VRAM (218 free) |
| A Swift-1.5 27B IQ3_S (Win) | on | 94% (47/50) | 10/10 | 19/20 | 18/20 | 51.2 | 210.1 | 6 | 14.6 | 1 | 12070 MiB VRAM (218 free) |
| B Bonsai-2-27B-Abl PQ2_0 (.54) | off | 72% (36/50) | 10/10 | 14/20 | 12/20 | 4.8 | 15.5 | 0 | 33.6 | 0 | 10135 MiB VRAM (2153 free) |
| B Bonsai-2-27B-Abl PQ2_0 (.54) | on | 94% (47/50) | 10/10 | 17/20 | 20/20 | 26.6 | 68.2 | 0 | 33.6 | 0 | 10135 MiB VRAM (2153 free) |
| C Qwen3.8-27B GSQ-RCO IQ3_S (Win) | off | 68% (34/50) | 10/10 | 16/20 | 8/20 | 9.4 | 55.6 | 0 | 14.8 | 2 | 12070 MiB VRAM (218 free) |
| C Qwen3.8-27B GSQ-RCO IQ3_S (Win) | on | 90% (45/50) | 10/10 | 19/20 | 16/20 | 59.9 | 207.4 | 9 | 14.9 | 5 | 12070 MiB VRAM (218 free) |
| D Gemma4-12B Q4_K_M+MTP (.54) | off | 66% (33/50) | 9/10 | 16/20 | 8/20 | 2.4 | 6.3 | 0 | 62.0 | 0 | 11187 MiB VRAM (1101 free) |
| D Gemma4-12B Q4_K_M+MTP (.54) | on | 76% (38/50) | 8/10 | 16/20 | 14/20 | 8.0 | 40.6 | 0 | 51.6 | 1 | 11187 MiB VRAM (1101 free) |
| E Xing4.0-29B-A4B IQ3_XXS (.54) | off | 36% (18/50) | 9/10 | 7/20 | 2/20 | 7.3 | 39.3 | 0 | 21.1 | 1 | 10921 MiB VRAM (1367 free) |
| E Xing4.0-29B-A4B IQ3_XXS (.54) | on | 32% (16/50) | 9/10 | 6/20 | 1/20 | 69.0 | 157.1 | 17 | 20.7 | 13 | 10921 MiB VRAM (1367 free) |
| F MiMo-9B MLX 4bit (Mac) | off | 52% (26/50) | 10/10 | 14/20 | 2/20 | 6.6 | 18.8 | 0 | 18.7 | 0 | ~7.2-9.4 GB unified (Mac) |
| F MiMo-9B MLX 4bit (Mac) | on | 52% (26/50) | 9/10 | 14/20 | 3/20 | 6.7 | 176.2 | 3 | 20.6 | 3 | ~7.2-9.4 GB unified (Mac) |
| F MiMo-9B MLX 4bit (Mac) | on_forced | 64% (32/50) | 9/10 | 16/20 | 7/20 | 17.4 | 162.9 | 2 | 18.9 | 2 | ~7.2-9.4 GB unified (Mac) |
| G Hermes-4-14B-OBL EXL3 4.0 (.54) | off | 8% (4/50) | 8/10 | 13/20 | 4/20 | 25.5 | 32.1 | 0 | 32.8 | 46 | 10669 MiB VRAM (1619 free; Q6 cache, chunk 512) |
| G Hermes-4-14B-OBL EXL3 4.0 (.54) | on | 10% (5/50) | 8/10 | 10/20 | 5/20 | 94.4 | 106.5 | 0 | 13.1 | 40 | 10669 MiB VRAM (1619 free; Q6 cache, chunk 512) |
| H Bonsai-2-27B MLX 2bit (Mac) | off | 62% (31/50) | 10/10 | 13/20 | 8/20 | 14.9 | 50.4 | 0 | 9.9 | 0 | 9966 MB footprint (Mac) |
| H Bonsai-2-27B MLX 2bit (Mac) | on | 94% (47/50) | 10/10 | 18/20 | 19/20 | 87.0 | 260.2 | 16 | 9.8 | 0 | 9966 MB footprint (Mac) |
| I Bonsai-2-27B-Abl PQ2_0 (Mac) | off | 72% (36/50) | 10/10 | 14/20 | 12/20 | 12.0 | 51.3 | 0 | 11.6 | 0 | ~10.0 GB footprint (Mac) |
| I Bonsai-2-27B-Abl PQ2_0 (Mac) | on | 90% (45/50) | 10/10 | 16/20 | 19/20 | 81.7 | 246.7 | 17 | 11.0 | 0 | ~10.0 GB footprint (Mac) |
| J54 Bonsai-2-27B PQ2_0+MTP (.54) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 3.8 | 9.7 | 0 | 53.0 | 0 | 10963 MiB VRAM (1325 free) |
| J54 Bonsai-2-27B PQ2_0+MTP (.54) | on | 96% (48/50) | 10/10 | 18/20 | 20/20 | 15.8 | 43.3 | 0 | 52.0 | 0 | 10963 MiB VRAM (1325 free) |
| JMac Bonsai-2-27B PQ2_0+MTP (Mac) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 13.1 | 50.8 | 0 | 12.1 | 0 | ~10.0 GB footprint (Mac) |
| JMac Bonsai-2-27B PQ2_0+MTP (Mac) | on | 96% (48/50) | 10/10 | 19/20 | 19/20 | 67.1 | 172.6 | 10 | 11.3 | 0 | ~10.0 GB footprint (Mac) |
| L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) | off | 68% (34/50) | 10/10 | 16/20 | 8/20 | 23.1 | 134.2 | 2 | 6.4 | 2 | 12.6 GB weights, footprint ~9.9 GB + paging (Mac) |
| L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) | on | 92% (46/50) | 10/10 | 18/20 | 18/20 | 81.9 | 310.9 | 14 | 6.2 | 0 | 12.6 GB weights, footprint ~9.9 GB + paging (Mac) |
| M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) | off | 46% (23/50) | 10/10 | 11/20 | 2/20 | 6.7 | 33.2 | 0 | 13.5 | 0 | 12 GB footprint peak (Mac) |
| M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) | on | 82% (41/50) | 10/10 | 17/20 | 14/20 | 65.0 | 227.1 | 10 | 13.3 | 1 | 12 GB footprint peak (Mac) |
| K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 8.5 | 31.5 | 0 | 14.8 | 0 | 11203 MiB VRAM (1085 free; draft on CPU) |
| K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) | on | 96% (48/50) | 10/10 | 19/20 | 19/20 | 52.0 | 243.7 | 7 | 14.1 | 1 | 11203 MiB VRAM (1085 free; draft on CPU) |
| KMac Bonsai-2 PQ2_0+DFlash2 (Mac) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 13.1 | 57.4 | 0 | 12.0 | 0 | ~9.3 GB weights, footprint ~8.5-10 GB (Mac) |
| KMac Bonsai-2 PQ2_0+DFlash2 (Mac) | on | 98% (49/50) | 10/10 | 19/20 | 20/20 | 71.6 | 172.2 | 11 | 11.6 | 0 | ~9.3 GB weights, footprint ~8.5-10 GB (Mac) |
| N54 HauhauCS Qwen3.8-27B IQ3_XS+MTP (.54, partial offload) | off | 60% (30/50) | 10/10 | 12/20 | 8/20 | 48.4 | 184.8 | 8 | 2.0 | 0 | 11221 MiB VRAM (1067 free; -ngl 44/65, rest on CPU) |
| NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) | off | 62% (31/50) | 10/10 | 13/20 | 8/20 | 17.6 | 60.7 | 0 | 7.6 | 0 | Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap) |
| NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) | on | 86% (43/50) | 10/10 | 18/20 | 15/20 | 135.5 | 467.6 | 27 | 6.8 | 5 | Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap) |
Notes: sampling for every request was temperature 0.2, top_p 0.95, top_k 20, seed 42. Thinking ON used max_tokens 3000; OFF used 800.
Requests ran one at a time. Wall time is measured at the client. Full per-task detail, repeat runs and caveats are in summary.md.
Winning configuration: exact J launch command (.54)
Binary: PrismML llama.cpp (prism-b10685 source) plus scripts/0001-qwen35-mtp-hadamard-inverse.patch, built with
scripts/build_prismml_mtp.sh (CUDA 12.8, sm_86). Model: Bonsai-2-27B-PQ2_0-MTP.gguf from
decent-jawfish/bonsai-2-27b-mtp (sha256 78df4279d40ebebdccfd2dae0e9d4847afee52e94f48f3542ae9437220dbd847).
As benchmarked (2026-09-27, -c 16384):
./llama-server -m Bonsai-2-27B-PQ2_0-MTP.gguf --alias bonsai-27b-pq2-mtp --host 0.0.0.0 --port 8801 \
-ngl 999 -fa on -c 16384 -np 1 -ctk q8_0 -ctv q8_0 --jinja \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--chat-template-kwargs '{"reasoning_effort": "medium"}' \
--spec-type draft-mtp --spec-draft-n-max 2
Production since 2026-09-28: same command with -c 28672 ... -ctkd q8_0 -ctvd q8_0 -fit off. That is the largest context that stays fully on the GPU with at least 300 MiB free (see below).
Max-context search on .54 (2026-09-28)
Conditions: fully on the GPU (-ngl 999, -fit off, no CPU offload), draft MTP on (--spec-draft-n-max 2), q8_0 KV for the target and the draft (-ctk/-ctv/-ctkd/-ctvd q8_0), -fa on, -np 1, with the vLLM embedder resident (2548 MiB).
A setting only counts if the server loads and a synthetic prompt filling about 90% of the context completes, with at least 300 MiB of GPU memory free at peak. The script is scripts/ctx_test.py and the raw results are in results/ctx_search_54.jsonl.
Each extra 2048 tokens of context costs about 46 MiB.
| Model | -c | Loads | 90% prompt (tokens) | Prompt tok/s | Decode tok/s @ 90% fill | Decode tok/s (short, JSON) | Draft accept | VRAM used, loaded (MiB, whole GPU) | Min free during prompt (MiB) | Counts (>=300 MiB free) |
|---|---|---|---|---|---|---|---|---|---|---|
| Abliterated-v2 PQ2_0 MTP | 16384 | OK | 14769 (90.1%) | 496.8 | 35.9 | 56.2 | 0.92 | 11059 | n/a (first run) | no, margin |
| Bonsai-2-27B PQ2_0 MTP | 16384 | OK | 14769 (90.1%) | 492.6 | 37.1 | 57.1 | 0.933 | 10963 | 933 | yes |
| Bonsai-2-27B PQ2_0 MTP | 24576 | OK | 22107 (90.0%) | 474.8 | 34.3 | 56.8 | 0.933 | 11333 | 563 | yes |
| Bonsai-2-27B PQ2_0 MTP | 28672 | OK | 25817 (90.0%) | 465.8 | 33.8 | 56.7 | 0.933 | 11517 | 379 | yes |
| Bonsai-2-27B PQ2_0 MTP | 30720 | OK | 27636 (90.0%) | 460.4 | 33.4 | 56.8 | 0.933 | 11609 | 287 | no, margin |
| Bonsai-2-27B PQ2_0 MTP | 32768 | OK | 29484 (90.0%) | 456.3 | 32.5 | 56.9 | 0.933 | 11701 | 195 | no, margin |
| Abliterated-v2 PQ2_0 MTP | 26624 | OK | 23977 (90.1%) | 470.4 | 33.6 | 56.0 | 0.92 | 11521 | 375 | yes |
| Abliterated-v2 PQ2_0 MTP | 28672 | OK | 25817 (90.0%) | 465.0 | 33.6 | 55.7 | 0.92 | 11613 | 283 | no, margin |
- Bonsai-2-27B PQ2_0 MTP: max safe
-c 28672(379 MiB free at peak). 30720 and 32768 also ran but left only 287 and 195 MiB free. - Abliterated-v2 PQ2_0 MTP: max safe
-c 26624(375 MiB free). The file is about 100 MB larger, so it fits one step less. - For comparison, the Windows 3060 has no embedder and reaches
-c 77824.
Abliterated variant with MTP
The bench’s “B” model was Ternary-Bonsai-2-27B-Abliterated-PQ2_0.gguf (no MTP). Kept on .54 for later use (not serving) is
BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF,
file Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf (sha256 a4e4c7b578131595c1694354bd6c74d00920df1f5082647a6126899753ebebf8, verified).
Its model card says v2 fixes v1’s thinking-mode blank answers. It loads with --spec-type draft-mtp on the same binary: 56 tok/s on a short JSON generation (draft acceptance 0.92) and 33.6 tok/s at 90% of a 26.6K context. It was not run through the 50-task suite.
Layout
code/:bench.py(runner),tasks.py/tasks_hard.py/tasks_ultra.py(prompts + graders),summarize.py,combined.py,final_table.py,vram_stats.py,speedtest.pyresults/: raw JSONL for every run (results_<X>.jsonl= easy,results_hard_<X>.jsonl,results_ultra_<X>.jsonl; letter = model key inbench.py), plusctx_search_54.jsonlsummary.md: full generated report (per-suite and per-task tables, repeat runs, caveats);final_table.md: the table abovescripts/: MTP patch, build script, and the context-search harness
To reproduce: cd code && LINUX54_HOST=<host> python3 bench.py --model J --modes off,on --suite tasks_hard --out results_hard_J.jsonl (--suite is tasks, tasks_hard or tasks_ultra; --base overrides the endpoint).
summarize.py and combined.py read the results*.jsonl files from the working directory.
Clean-up after the bench (2026-09-28)
Only Bonsai models were kept (base MTP on .54, Windows and Mac, plus the abliterated v2 MTP on .54). The Swift/GSQ/Gemma/Xing/MiMo/Hermes/ThinkingCap/ZDTaichu/HauhauCS weights were deleted.