Plain-language summary of what won and why. Full data and methodology: README.md (rounds 1-8), data/, scripts/ in this folder.
The winners
Hard reasoning tasks: thinking mode at temperature 1.0 with presence penalty 1.0, sharp template. It solved the hard T3 decode task in 4 of 7 tries (57%), produced the best test cell ever measured here (quality 9.3 in 135 seconds), and did it with about 3x less thinking text than the untuned settings.
Quick answers: instruct mode with the official params (temperature 0.7, presence 1.5). It crushes the easy tasks in seconds with no thinking overhead at all.
The model: the v2 quant. Same 4.27 bpw bitrate as the old Q4_K_M, but it is the only Q4_K_M that ever solved T3, and it matches or beats the bigger UD-IQ4_XS.
Does it make sense? Yes.
The v2 quant winning makes sense because its importance matrix was built from the right data this time: the actual abliterated model generating real thinking traces on agentic coding content, instead of a stock model’s activations. A quantizer that knows what the model actually does makes better compression decisions. Same recipe, better ingredients.
The presence penalty winning is the most interesting result, and it also makes sense. T3’s failure mode was never “too dumb to solve it” — it was getting stuck in a loop, re-deriving the same decode steps until the 16,384-token cap burned out with no answer. Presence penalty penalizes repetition, which is exactly the medicine for a loop. The evidence fits: the wins with presence 1.0 are faster and shorter, and one captured run had found the flag in its reasoning but kept rambling past the cap. The reasoning ability was always there; the off-switch was broken. Presence fixes the off-switch. That is also why greedy fails (a deterministic loop never escapes) and why raising the cap to 24K did not help (you cannot out-wait a loop).
Instruct winning the easy tasks is trivially true — thinking is pure cost when the answer is easy. A 60K-character overthink on a one-line Suricata rule task proves it: thinking mode produced zero final content there, while instruct answered in 44 seconds.
Caveats that keep this honest
- Presence 1.0 at 4/7 versus 2/7 for the official setting is suggestive, not statistically proven — the samples are small.
- Presence 1.0 trades one risk for another: occasionally it stops too early with a wrong answer instead of rambling forever. It is task-specific medicine for loop-prone problems, not a universal improvement over Qwen’s official settings.
- Single-shot T3 scores are a lottery; only pass rates over several runs mean anything.
- Sharp over baked matches the earlier template experiments in this repo, which is what a real result should look like.
What is deployed
Router :12434, family qwen38-flash-next-unc-q4km-agentic-v2:
- thinking (default via the bare alias qwen38-unc-q4km-v2): temperature 1.0, top_p 0.95, top_k 20, min_p 0.0, presence 1.0 — the evidence-tuned winner.
- instruct: official 0.7 / 0.80 / 20 / presence 1.5 — for speed, not hard reasoning. Backend :12702 serves the same defaults with sharp template, 225,280 ctx, mmproj, Q8 KV. The MTP draft head stays enabled: a 5-task A/B on this quant (round 9) showed +48-59% t/s on long generations with unchanged quality — the older “MTP is only for greedy” conclusion did not replicate on v2.