Back to the experimentSupporting notebook

Unc Q4_K_M agentic-v2: fixed imatrix, requant, and sampling findings

Source: qwen38/flash-next-unc-q4km-agentic-v2-2026-09-08/README.md · revision 6550ead3945b

Plain-language summary: FINDINGS.md — winners, why they make sense, caveats, deployed config.

This experiment fixes the two defects of the prior quant job (quant-jobs/fp8-unc-udiq4xs-6h): a live importance matrix that timed out at ~85% and was backfilled with an AtomicChat stock imatrix (wrong activation distribution for the abliterated model), and calibration data with no real long <think> traces. It then re-quantizes Q4_K_M with the identical tensor map, benches the same auto-graded cybersecurity suite (T1 log4shell / T2 suricata / T3 CTF-decode, 16,384-token cap unless noted), and — the main finding — establishes sampling methodology for the hard T3 task.

1. Fixed imatrix: two complete 150-chunk passes, merged only with each other

Step Detail
Calibration 570 agentic-coding docs (~920K chars) + 100 real thinking traces generated from the ablit model itself (reasoning ≥ 400 chars, median 1,912; data/thinking_traces_calib.jsonl). Total 670 docs ≈ 374K tokens, two independent shuffles (scripts/build_calib.py, scripts/gen_traces.py)
Pass A --chunks 150 -c 1024 --parse-special --no-ppl, all 3 GPUs (-ngl 99 -sm layer -ts 1,1,1, PLE + 24/48 expert blocks on CPU), completed rc=0 in ~95 min (~38 s/chunk steady)
Pass B Same, completed rc=0 in ~101 min (comma-separated single --override-tensor form, zero deprecation warnings)
Merge llama-imatrix -m MASTER --in-file pass_a,pass_b — no AtomicChat data anywhere (-m is required by arg validation; no model load happens in merge mode). See scripts/run_imatrix.sh
Coverage Pass A final save: experts min 95.3% / avg 99.4%. Pass B: min 98.6% / avg 99.7%. Merged save: no partial-data warnings. Sub-100% per-pass expert coverage is inherent to MoE routing, not a defect

2. Requant: same AD-4.27bpw Q4_K_M-M64 tensor map

llama-quantize --imatrix merged.agentic --allow-requantize --tensor-type-file AD-4.27bpw-Q4_K_M-M64.tensor-types (835 overrides, 145 applied — identical counts to the prior Q4_K_M). 36 min + split. Staged at WORKSTATION/AD-4.27bpw-Q4_K_M-M64-agentic-v2.

3. Bench results (4-cell suite, data/bench_results_v2.json)

cell stock Q4_K_M (old) ablit IQ4_XS (old) v2 agentic-imatrix
thinking | sharp q=6.0 / 260 s q=9.3 / 365 s q=9.3 / 346 s
thinking | baked q=6.0 / 700 s q=6.0 / 220 s q=9.3 / 253 s
instruct | sharp q=6.0 / 159 s q=6.0 / 205 s q=6.0 / 126 s
instruct | baked q=5.0 / 308 s q=6.0 / 230 s q=6.0 / 198 s

v2 is the only Q4_K_M that solves T3 (both templates), matches the ablit IQ4_XS best case, and uses the least VRAM (61.1 GiB vs 63.3/67.7). T1/T2 throughput 50–59 t/s, no regression.

4. Main finding: T3 is a sampling lottery — presence_penalty is the lever

Single-shot T3 quality is meaningless: the same config passes or fails run to run (bimodal — converges in ~6–12K tokens or rambles to the 16,384 cap with zero content). Measured pass rates (all thinking|sharp unless noted):

sampling T3 pass rate evidence
temp 0.6, presence 0.0 2/5 data/t3_passrate_220k_mmproj.json
temp 1.0 (official), presence 0.0 1/5 data/t3_passrate_official_temp10.json
greedy (temp 0) 0/1 round-2 probe
instruct official (0.7/0.80/20 + presence 1.5) 1/8 data/t3_passrate_router_instruct*.json
thinking 1.0 + presence 0.5 1/2 data/t3_thinking_presence_probe.json
thinking 1.0 + presence 1.0 3/5 same probe + confirm runs
thinking 1.0 + presence 1.0, max_tokens 24K 1/2 data/t3_pres10_24k.json — raising the cap does not convert failures

Best measured cell: thinking + presence 1.0 | sharp → q=9.3 in 134.7 s (data/bench_results_v2_thinkpres10_sharp.json), T3 solved in 81.5 s with only 7.1K reasoning chars — presence 1.0 compresses reasoning ~2–4x vs presence 0.0 (15–30K chars).

Supporting observations:

5. Production configuration adopted

Router :12434 family qwen38-flash-next-unc-q4km-agentic-v2 (forced presets, verified live):

Backend :12702: 225,280 ctx, mmproj F16, KV q8_0, sharp template, server defaults aligned to thinking 1.0/0.95/20/presence 1.0 (scripts/run_v2_220k_mmproj.sh).

6. Honest caveats

Charts

Generated by scripts/make_charts.py from data/*.json.

Round 8 — T4/T5 extension + 5-task suite (2026-09-08)

Round 9 — MTP draft A/B on v2: the draft stays (2026-09-09)

The MTP experiment (flash-next-unc-mtp-draft-2026-09-07) concluded “keep the plain server” (parity at temp 0.7, hostile to spec decode at prod sampling) — yet the v2 backend shipped WITH the draft. Before “fixing” it, we A/B’d on the 5-task suite, production profiles only, single variable = -md draft (data/bench_5task_draftON.json / bench_5task_draftOFF.json, scripts/bench_5task_ab.py):

profile T3 t/s draftON T3 t/s draftOFF quality
thinking-pres1.0 39.9 27.0 (+48% draft) T1/T2/T4/T5 all 9/9 both arms; T3 lottery (ON fail / OFF pass — single-shot, not evidence)
instruct-official 46.0 29.0 (+59% draft) same pattern; T3 fail both

Conclusion: on v2 the draft is a clear throughput win at production sampling — the 09-07 “parity” conclusion does not replicate here (it carried a ±20% CPU-contention caveat and ran on the older agentic quant). Draft restored to production; service verified healthy post-restart. Load average ~4 during both arms; short tasks (<10 s) are too noisy to compare — the long T3 generations are the reliable signal.

Reproduction

  1. Calib: scripts/gen_traces.py <port> against any ablit Unc backend, then scripts/build_calib.py.
  2. Imatrix: scripts/run_imatrix.sh (phases A/B, then merge with -m; ~3.5 h on 3x3090 for 2×150 chunks).
  3. Quantize: prior-job quantize_orcarouter_fp8.sh with SKIP_CONVERT=1, same map, merged imatrix.
  4. Serve: scripts/run_v2_220k_mmproj.sh (needs the v2 weights + mmproj).
  5. Bench: scripts/bench_official_12702.py (cells), scripts/t3_passrate_temp10.py / scripts/t3_thinking_presence.py / scripts/t3_passrate_router_instruct.py (T3 pass rates).