Plain-language summary: FINDINGS.md — winners, why they make sense, caveats, deployed config.
- Date: 2026-09-06 → 2026-09-08 (quant job + 6 bench rounds)
- Host:
12700— 3x RTX 3090 (72 GB aggregate VRAM), 20 CPUs, Ubuntu Linux - Model:
orcarouter/Qwen3.8-Flash-Next-Uncensored-FP8(abliterated “Heretic” 125B-A6B MoE) → q8_0 master (188 GB, reused) → AD-4.27bpw Q4_K_M-M64, 89 GB single file (shards consolidated post-quant) - Inference stack: llama.cpp flash-next fork (
--repack --op-offload,-sm layer -ts 1,1,1), Q8_0 KV cache, 225,280-token context, mmproj F16 active, sharp chat template - Router:
qwen36-multi-variant-proxyon:12434, familyqwen38-flash-next-unc-q4km-agentic-v2(qwen38-unc-q4km-instruct/-thinking), backend:12702
This experiment fixes the two defects of the prior quant job (quant-jobs/fp8-unc-udiq4xs-6h):
a live importance matrix that timed out at ~85% and was backfilled with an AtomicChat stock
imatrix (wrong activation distribution for the abliterated model), and calibration data with no
real long <think> traces. It then re-quantizes Q4_K_M with the identical tensor map, benches
the same auto-graded cybersecurity suite (T1 log4shell / T2 suricata / T3 CTF-decode, 16,384-token
cap unless noted), and — the main finding — establishes sampling methodology for the hard T3 task.
1. Fixed imatrix: two complete 150-chunk passes, merged only with each other
| Step | Detail |
|---|---|
| Calibration | 570 agentic-coding docs (~920K chars) + 100 real thinking traces generated from the ablit model itself (reasoning ≥ 400 chars, median 1,912; data/thinking_traces_calib.jsonl). Total 670 docs ≈ 374K tokens, two independent shuffles (scripts/build_calib.py, scripts/gen_traces.py) |
| Pass A | --chunks 150 -c 1024 --parse-special --no-ppl, all 3 GPUs (-ngl 99 -sm layer -ts 1,1,1, PLE + 24/48 expert blocks on CPU), completed rc=0 in ~95 min (~38 s/chunk steady) |
| Pass B | Same, completed rc=0 in ~101 min (comma-separated single --override-tensor form, zero deprecation warnings) |
| Merge | llama-imatrix -m MASTER --in-file pass_a,pass_b — no AtomicChat data anywhere (-m is required by arg validation; no model load happens in merge mode). See scripts/run_imatrix.sh |
| Coverage | Pass A final save: experts min 95.3% / avg 99.4%. Pass B: min 98.6% / avg 99.7%. Merged save: no partial-data warnings. Sub-100% per-pass expert coverage is inherent to MoE routing, not a defect |
2. Requant: same AD-4.27bpw Q4_K_M-M64 tensor map
llama-quantize --imatrix merged.agentic --allow-requantize --tensor-type-file AD-4.27bpw-Q4_K_M-M64.tensor-types
(835 overrides, 145 applied — identical counts to the prior Q4_K_M). 36 min + split.
Staged at WORKSTATION/AD-4.27bpw-Q4_K_M-M64-agentic-v2.
3. Bench results (4-cell suite, data/bench_results_v2.json)
| cell | stock Q4_K_M (old) | ablit IQ4_XS (old) | v2 agentic-imatrix |
|---|---|---|---|
| thinking | sharp | q=6.0 / 260 s | q=9.3 / 365 s | q=9.3 / 346 s |
| thinking | baked | q=6.0 / 700 s | q=6.0 / 220 s | q=9.3 / 253 s |
| instruct | sharp | q=6.0 / 159 s | q=6.0 / 205 s | q=6.0 / 126 s |
| instruct | baked | q=5.0 / 308 s | q=6.0 / 230 s | q=6.0 / 198 s |
v2 is the only Q4_K_M that solves T3 (both templates), matches the ablit IQ4_XS best case, and uses the least VRAM (61.1 GiB vs 63.3/67.7). T1/T2 throughput 50–59 t/s, no regression.
4. Main finding: T3 is a sampling lottery — presence_penalty is the lever
Single-shot T3 quality is meaningless: the same config passes or fails run to run (bimodal — converges in ~6–12K tokens or rambles to the 16,384 cap with zero content). Measured pass rates (all thinking|sharp unless noted):
| sampling | T3 pass rate | evidence |
|---|---|---|
| temp 0.6, presence 0.0 | 2/5 | data/t3_passrate_220k_mmproj.json |
| temp 1.0 (official), presence 0.0 | 1/5 | data/t3_passrate_official_temp10.json |
| greedy (temp 0) | 0/1 | round-2 probe |
| instruct official (0.7/0.80/20 + presence 1.5) | 1/8 | data/t3_passrate_router_instruct*.json |
| thinking 1.0 + presence 0.5 | 1/2 | data/t3_thinking_presence_probe.json |
| thinking 1.0 + presence 1.0 | 3/5 | same probe + confirm runs |
| thinking 1.0 + presence 1.0, max_tokens 24K | 1/2 | data/t3_pres10_24k.json — raising the cap does not convert failures |
Best measured cell: thinking + presence 1.0 | sharp → q=9.3 in 134.7 s
(data/bench_results_v2_thinkpres10_sharp.json), T3 solved in 81.5 s with only 7.1K
reasoning chars — presence 1.0 compresses reasoning ~2–4x vs presence 0.0 (15–30K chars).
Supporting observations:
- One presence-1.0 run burned the full cap yet contained the flag in its reasoning trace (found it, kept rambling, never answered) — a stop-discipline failure, not reasoning failure.
- Higher temperature increases variance (1.0 → 20% vs 0.6 → 40% at presence 0.0).
- Sharp template beats baked consistently (baked officials: both q=6.0, T3 fail).
- Long thinking that produces correct answers is cost, not waste; genuine waste is long thinking with no answer (e.g. stock Q4_K_M: 52K reasoning chars on trivial T2, zero content).
5. Production configuration adopted
Router :12434 family qwen38-flash-next-unc-q4km-agentic-v2 (forced presets, verified live):
- thinking: temp 1.0, top_p 0.95, top_k 20, min_p 0.0, presence 1.0 (evidence-tuned; deliberate deviation from official 0.0), repeat 1.0, thinking+preserve on.
- instruct: official 0.7 / 0.80 / 20 / min_p 0.0 / presence 1.5 / repeat 1.0, thinking off (best-tested non-thinking config; for speed, not hard reasoning).
Backend :12702: 225,280 ctx, mmproj F16, KV q8_0, sharp template, server defaults aligned
to thinking 1.0/0.95/20/presence 1.0 (scripts/run_v2_220k_mmproj.sh).
6. Honest caveats
- T3 samples are small (n=5 per config); 3/5 vs 2/5 is suggestive, not statistically settled.
- The abliteration’s longer-thinking trait is unchanged by the imatrix fix; what changed is that the thinking budget now converts into correct answers on this suite.
- Per-pass expert coverage remains ~95–99% minima; the merge of two passes mitigates it.
Charts
Generated by scripts/make_charts.py from data/*.json.
Round 8 — T4/T5 extension + 5-task suite (2026-09-08)
- T4 (YARA rule, medium) and T5 (SSH-log forensics, hard-but-distinct) added
(
scripts/tasks_t4t5.py): both score 9/9 on all 3 profiles, all regex checks green — well-calibrated, non-discriminating tasks. - 5-task cells (
data/bench_results_v2_5task.json): thinking+pres1.0 q=7.2, thinking-official q=9.2, instruct-official q=7.2. T3 is the only discriminator (pres1.0 fail / official pass / instruct fail — single-shot variance again). - Updated T3 aggregates: think1.0+pres1.0 → 4/7 (57%); think1.0+pres0 → 2/7 (29%); instruct-official → 1/8; think0.6/pres0 → 2/5. Presence 1.0 still leads.
- T2 pathology: both thinking runs burned ~60K reasoning chars with zero content on the trivial T2 rule task (passed via reasoning trace only) — presence 1.0 does not prevent runaway overthinking on easy tasks. Instruct did T2 in 44 s / 320 chars.
Round 9 — MTP draft A/B on v2: the draft stays (2026-09-09)
The MTP experiment (flash-next-unc-mtp-draft-2026-09-07) concluded “keep the plain server”
(parity at temp 0.7, hostile to spec decode at prod sampling) — yet the v2 backend shipped
WITH the draft. Before “fixing” it, we A/B’d on the 5-task suite, production profiles only,
single variable = -md draft (data/bench_5task_draftON.json / bench_5task_draftOFF.json,
scripts/bench_5task_ab.py):
| profile | T3 t/s draftON | T3 t/s draftOFF | quality |
|---|---|---|---|
| thinking-pres1.0 | 39.9 | 27.0 (+48% draft) | T1/T2/T4/T5 all 9/9 both arms; T3 lottery (ON fail / OFF pass — single-shot, not evidence) |
| instruct-official | 46.0 | 29.0 (+59% draft) | same pattern; T3 fail both |
Conclusion: on v2 the draft is a clear throughput win at production sampling — the 09-07 “parity” conclusion does not replicate here (it carried a ±20% CPU-contention caveat and ran on the older agentic quant). Draft restored to production; service verified healthy post-restart. Load average ~4 during both arms; short tasks (<10 s) are too noisy to compare — the long T3 generations are the reliable signal.
Reproduction
- Calib:
scripts/gen_traces.py <port>against any ablit Unc backend, thenscripts/build_calib.py. - Imatrix:
scripts/run_imatrix.sh(phases A/B, then merge with-m; ~3.5 h on 3x3090 for 2×150 chunks). - Quantize: prior-job
quantize_orcarouter_fp8.shwithSKIP_CONVERT=1, same map, merged imatrix. - Serve:
scripts/run_v2_220k_mmproj.sh(needs the v2 weights + mmproj). - Bench:
scripts/bench_official_12702.py(cells),scripts/t3_passrate_temp10.py/scripts/t3_thinking_presence.py/scripts/t3_passrate_router_instruct.py(T3 pass rates).