Journal

Rebuilding calibration for an agentic quant

The clean agentic importance matrix recovered the hard T3 task; later draft decoding improved measured throughput.

Vlad / experimentos.
The boundary

The hard decode task is stochastic. Small pass-rate probes do not establish a general quality ranking.

The question

Calibration was rebuilt around agentic work rather than assumed adequate from a generic matrix. The hard task, new quantization artifacts and subsequent draft-decoding probes stayed separate evidence gates.

From the original notebook

Plain-language summary: FINDINGS.md — winners, why they make sense, caveats, deployed config.

  • Date: 2026-09-06 → 2026-09-08 (quant job + 6 bench rounds)
  • Host: 12700 — 3x RTX 3090 (72 GB aggregate VRAM), 20 CPUs, Ubuntu Linux
  • Model: orcarouter/Qwen3.8-Flash-Next-Uncensored-FP8 (abliterated “Heretic” 125B-A6B MoE) → q8_0 master (188 GB, reused) → AD-4.27bpw Q4_K_M-M64, 89 GB single file (shards consolidated post-quant)
  • Inference stack: llama.cpp flash-next fork (--repack --op-offload, -sm layer -ts 1,1,1), Q8_0 KV cache, 225,280-token context, mmproj F16 active, sharp chat template
  • Router: qwen36-multi-variant-proxy on :12434, family qwen38-flash-next-unc-q4km-agentic-v2 (qwen38-unc-q4km-instruct / -thinking), backend :12702

This experiment fixes the two defects of the prior quant job (quant-jobs/fp8-unc-udiq4xs-6h): a live importance matrix that timed out at ~85% and was backfilled with an AtomicChat stock imatrix (wrong activation distribution for the abliterated model), and calibration data with no real long <think> traces. It then re-quantizes Q4_K_M with the identical tensor map, benches the same auto-graded cybersecurity suite (T1 log4shell / T2 suricata / T3 CTF-decode, 16,384-token cap unless noted), and — the main finding — establishes sampling methodology for the hard T3 task.

1. Fixed imatrix: two complete 150-chunk passes, merged only with each other

Step Detail
Calibration 570 agentic-coding docs (~920K chars) + 100 real thinking traces generated from the ablit model itself (reasoning ≥ 400 chars, median 1,912; data/thinking_traces_calib.jsonl). Total 670 docs ≈ 374K tokens, two independent shuffles (scripts/build_calib.py, scripts/gen_traces.py)
Pass A --chunks 150 -c 1024 --parse-special --no-ppl, all 3 GPUs (-ngl 99 -sm layer -ts 1,1,1, PLE + 24/48 expert blocks on CPU), completed rc=0 in ~95 min (~38 s/chunk steady)
Pass B Same, completed rc=0 in ~101 min (comma-separated single --override-tensor form, zero deprecation warnings)
Merge llama-imatrix -m MASTER --in-file pass_a,pass_b — no AtomicChat data anywhere (-m is required by arg validation; no model load happens in merge mode). See scripts/run_imatrix.sh
Coverage Pass A final save: experts min 95.3% / avg 99.4%. Pass B: min 98.6% / avg 99.7%. Merged save: no partial-data warnings. Sub-100% per-pass expert coverage is inherent to MoE routing, not a defect

2. Requant: same AD-4.27bpw Q4_K_M-M64 tensor map

llama-quantize --imatrix merged.agentic --allow-requantize --tensor-type-file AD-4.27bpw-Q4_K_M-M64.tensor-types (835 overrides, 145 applied — identical counts to the prior Q4_K_M). 36 min + split. Staged at WORKSTATION/AD-4.27bpw-Q4_K_M-M64-agentic-v2.

3. Bench results (4-cell suite, data/bench_results_v2.json)

cell stock Q4_K_M (old) ablit IQ4_XS (old) v2 agentic-imatrix
thinking | sharp q=6.0 / 260 s q=9.3 / 365 s q=9.3 / 346 s
thinking | baked q=6.0 / 700 s q=6.0 / 220 s q=9.3 / 253 s
instruct | sharp q=6.0 / 159 s q=6.0 / 205 s q=6.0 / 126 s
instruct | baked q=5.0 / 308 s q=6.0 / 230 s q=6.0 / 198 s

v2 is the only Q4_K_M that solves T3 (both templates), matches the ablit IQ4_XS best case, and uses the least VRAM (61.1 GiB vs 63.3/67.7). T1/T2 throughput 50–59 t/s, no regression.

4. Main finding: T3 is a sampling lottery — presence_penalty is the lever

Single-shot T3 quality is meaningless: the same config passes or fails run to run (bimodal — converges in ~6–12K tokens or rambles to the 16,384 cap with zero content). Measured pass rates (all thinking|sharp unless noted):

sampling T3 pass rate evidence
temp 0.6, presence 0.0 2/5 data/t3_passrate_220k_mmproj.json
temp 1.0 (official), presence 0.0 1/5 data/t3_passrate_official_temp10.json
greedy (temp 0) 0/1 round-2 probe
instruct official (0.7/0.80/20 + presence 1.5) 1/8 data/t3_passrate_router_instruct*.json
thinking 1.0 + presence 0.5 1/2 data/t3_thinking_presence_probe.json
thinking 1.0 + presence 1.0 3/5 same probe + confirm runs
thinking 1.0 + presence 1.0, max_tokens 24K 1/2 data/t3_pres10_24k.json — raising the cap does not convert failures

Best measured cell: thinking + presence 1.0 | sharp → q=9.3 in 134.7 s (data/bench_results_v2_thinkpres10_sharp.json), T3 solved in 81.5 s with only 7.1K reasoning chars — presence 1.0 compresses reasoning ~2–4x vs presence 0.0 (15–30K chars).

Supporting observations:

  • One presence-1.0 run burned the full cap yet contained the flag in its reasoning trace (found it, kept rambling, never answered) — a stop-discipline failure, not reasoning failure.
  • Higher temperature increases variance (1.0 → 20% vs 0.6 → 40% at presence 0.0).
  • Sharp template beats baked consistently (baked officials: both q=6.0, T3 fail).
  • Long thinking that produces correct answers is cost, not waste; genuine waste is long thinking with no answer (e.g. stock Q4_K_M: 52K reasoning chars on trivial T2, zero content).

5. Production configuration adopted

Router :12434 family qwen38-flash-next-unc-q4km-agentic-v2 (forced presets, verified live):

  • thinking: temp 1.0, top_p 0.95, top_k 20, min_p 0.0, presence 1.0 (evidence-tuned; deliberate deviation from official 0.0), repeat 1.0, thinking+preserve on.
  • instruct: official 0.7 / 0.80 / 20 / min_p 0.0 / presence 1.5 / repeat 1.0, thinking off (best-tested non-thinking config; for speed, not hard reasoning).

Backend :12702: 225,280 ctx, mmproj F16, KV q8_0, sharp template, server defaults aligned to thinking 1.0/0.95/20/presence 1.0 (scripts/run_v2_220k_mmproj.sh).

6. Honest caveats

  • T3 samples are small (n=5 per config); 3/5 vs 2/5 is suggestive, not statistically settled.
  • The abliteration’s longer-thinking trait is unchanged by the imatrix fix; what changed is that the thinking budget now converts into correct answers on this suite.
  • Per-pass expert coverage remains ~95–99% minima; the merge of two passes mitigates it.

Charts

Generated by scripts/make_charts.py from data/*.json.

Round 8 — T4/T5 extension + 5-task suite (2026-09-08)

  • T4 (YARA rule, medium) and T5 (SSH-log forensics, hard-but-distinct) added (scripts/tasks_t4t5.py): both score 9/9 on all 3 profiles, all regex checks green — well-calibrated, non-discriminating tasks.
  • 5-task cells (data/bench_results_v2_5task.json): thinking+pres1.0 q=7.2, thinking-official q=9.2, instruct-official q=7.2. T3 is the only discriminator (pres1.0 fail / official pass / instruct fail — single-shot variance again).
  • Updated T3 aggregates: think1.0+pres1.0 → 4/7 (57%); think1.0+pres0 → 2/7 (29%); instruct-official → 1/8; think0.6/pres0 → 2/5. Presence 1.0 still leads.
  • T2 pathology: both thinking runs burned ~60K reasoning chars with zero content on the trivial T2 rule task (passed via reasoning trace only) — presence 1.0 does not prevent runaway overthinking on easy tasks. Instruct did T2 in 44 s / 320 chars.

Round 9 — MTP draft A/B on v2: the draft stays (2026-09-09)

The MTP experiment (flash-next-unc-mtp-draft-2026-09-07) concluded “keep the plain server” (parity at temp 0.7, hostile to spec decode at prod sampling) — yet the v2 backend shipped WITH the draft. Before “fixing” it, we A/B’d on the 5-task suite, production profiles only, single variable = -md draft (data/bench_5task_draftON.json / bench_5task_draftOFF.json, scripts/bench_5task_ab.py):

profile T3 t/s draftON T3 t/s draftOFF quality
thinking-pres1.0 39.9 27.0 (+48% draft) T1/T2/T4/T5 all 9/9 both arms; T3 lottery (ON fail / OFF pass — single-shot, not evidence)
instruct-official 46.0 29.0 (+59% draft) same pattern; T3 fail both

Conclusion: on v2 the draft is a clear throughput win at production sampling — the 09-07 “parity” conclusion does not replicate here (it carried a ±20% CPU-contention caveat and ran on the older agentic quant). Draft restored to production; service verified healthy post-restart. Load average ~4 during both arms; short tasks (<10 s) are too noisy to compare — the long T3 generations are the reliable signal.

Reproduction

  1. Calib: scripts/gen_traces.py <port> against any ablit Unc backend, then scripts/build_calib.py.
  2. Imatrix: scripts/run_imatrix.sh (phases A/B, then merge with -m; ~3.5 h on 3x3090 for 2×150 chunks).
  3. Quantize: prior-job quantize_orcarouter_fp8.sh with SKIP_CONVERT=1, same map, merged imatrix.
  4. Serve: scripts/run_v2_220k_mmproj.sh (needs the v2 weights + mmproj).
  5. Bench: scripts/bench_official_12702.py (cells), scripts/t3_passrate_temp10.py / scripts/t3_thinking_presence.py / scripts/t3_passrate_router_instruct.py (T3 pass rates).

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS