Back to the experimentSupporting notebook

Findings in simple terms

Source: qwen38/flash-next-unc-q4km-agentic-v2-2026-09-08/FINDINGS.md · revision 6550ead3945b

Plain-language summary of what won and why. Full data and methodology: README.md (rounds 1-8), data/, scripts/ in this folder.

The winners

Hard reasoning tasks: thinking mode at temperature 1.0 with presence penalty 1.0, sharp template. It solved the hard T3 decode task in 4 of 7 tries (57%), produced the best test cell ever measured here (quality 9.3 in 135 seconds), and did it with about 3x less thinking text than the untuned settings.

Quick answers: instruct mode with the official params (temperature 0.7, presence 1.5). It crushes the easy tasks in seconds with no thinking overhead at all.

The model: the v2 quant. Same 4.27 bpw bitrate as the old Q4_K_M, but it is the only Q4_K_M that ever solved T3, and it matches or beats the bigger UD-IQ4_XS.

Does it make sense? Yes.

The v2 quant winning makes sense because its importance matrix was built from the right data this time: the actual abliterated model generating real thinking traces on agentic coding content, instead of a stock model’s activations. A quantizer that knows what the model actually does makes better compression decisions. Same recipe, better ingredients.

The presence penalty winning is the most interesting result, and it also makes sense. T3’s failure mode was never “too dumb to solve it” — it was getting stuck in a loop, re-deriving the same decode steps until the 16,384-token cap burned out with no answer. Presence penalty penalizes repetition, which is exactly the medicine for a loop. The evidence fits: the wins with presence 1.0 are faster and shorter, and one captured run had found the flag in its reasoning but kept rambling past the cap. The reasoning ability was always there; the off-switch was broken. Presence fixes the off-switch. That is also why greedy fails (a deterministic loop never escapes) and why raising the cap to 24K did not help (you cannot out-wait a loop).

Instruct winning the easy tasks is trivially true — thinking is pure cost when the answer is easy. A 60K-character overthink on a one-line Suricata rule task proves it: thinking mode produced zero final content there, while instruct answered in 44 seconds.

Caveats that keep this honest

What is deployed

Router :12434, family qwen38-flash-next-unc-q4km-agentic-v2: