Back to the experimentSupporting notebook

Qwen3.8-Flash-Next

Source: qwen-next-flash/README.md · revision 6550ead3945b

Also see the in-progress official-BF16 EXL3 campaign: exl3-350hq-375-2026-09-15 (3.50 HQ + 3.75 MTP6, plan + journey).

udq4kxl-cpu-moe-200k-2026-09-20 — UD-Q4_K_XL GGUF (104 GiB) CPU-MoE deployment: -ot expert placement tuning, 12/12/12 GPU split, 31.1 tok/s @ 204800/q8.

vllm-int4-autoround-3x3090-2026-09-20 — INT4-Mixed-AutoRound via patched vLLM PP3 podman deploy: 67.4 tok/s 1-stream / 73 aggregate @ 262144. New champion.

Qwen3.8-Flash-Next abliterated UD-IQ4_XS: quantization, deployment and template A/B

This experiment covers the complete pipeline for moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 → Unsloth-style UD-IQ4_XS GGUF: torchao-FP8 conversion, live agentic importance-matrix collection, proxy family deployment as the hermes agent default, and an 8-cell benchmark (2 model variants × thinking on/off × chat template sharp/baked) on auto-graded cybersecurity tasks at a 16,384-token completion cap.

1. Quantization job (quant-jobs/fp8-unc-udiq4xs-6h)

Step Detail
Source moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 (abliterated 125B A6B MoE, native block-FP8 via torchao, 512 experts/layer, PLE n-gram table 320M rows × 160 dims)
Converter patches torchao layout handling: per-expert sharded .weight_scale and per-layer scales for the PLE table (mandatory, otherwise scale lookup fails)
Master q8_0 re-quantization of all 1224 tensors → 188.2 GB single-file master
Importance matrix 170/200 chunks of live agentic traffic (2 h timebox), collected with CPU-MoE expert split (-ts 1,1,1; 16 GPU layer-blocks / 32 CPU layer-blocks, PLE on CPU), merged with an AtomicChat-calibration imatrix (4,000 chunks) → 579 MB file, 99.8% expert coverage at collection end
Quantization llama-quantize with the Unsloth UD-IQ4_XS tensor-type map (643 overrides; experts mostly iq3_s, PLE table iq4_nl, attention Q/K/MMLP at higher precision)
Output 88 GB, 41 shards → WORKSTATION/Qwen38FN-Unc-FP8-unsloth-UD-IQ4_XS-00001-of-00041.gguf
Smoke test (stock build, CPU experts) prompt 75.6 t/s, generation 21.8 t/s, correct reply

2. Proxy deployment (clone of the UD-Q4 family pattern)

Serving parameters were cloned from the qwen38-unsloth-q4 (UD-Q4_K_XL) family: sharp chat-template hook, Q8_0 KV cache, 204,800-token context, batch 16,384 / ubatch 400, 14 pinned CPU threads (0-13), metrics enabled. Hardware placement follows the existing flash-next families because the model (88 GB) exceeds VRAM (72 GB):

Root cause found during bring-up

Requests through the proxy hung indefinitely (clients saw HTTP 000/timeouts) while the backend itself was healthy. _backend_is_ready() in the proxy requires /v1/models to advertise exactly the family’s backend_model string. Our launch script used --alias qwen38abq4-200k while the family key is qwen38-flash-next-ablit-udiq4xs-200k → never “ready” → the GPU-acquisition wait (900 s) deadlocked on VRAM held by the proxy’s own backend. Fix: SERVED_MODEL_NAME must equal backend_model (the convention every other family follows). After the fix, a cold proxy launch from warm page cache answers its first request in ~17 s.

Standalone verification (before moving to the proxy)

A tmux-deployed standalone server on 0.0.0.0:12703 with a watchdog (hard deadline 1,500 s, 10 s progress trail, automatic self-test) was used to validate behaviour in isolation: ready in 21 s from warm page cache, self-test reply correct, 67.7 GB VRAM, 86.7 t/s single-stream non-thinking generation. See scripts/abq4_standalone.sh.

Known limitation: vision

The flash-next fork’s mtmd loader aborts (ggml_backend_buffer_set_usage in clip_model_loader::load_tensors) when loading this model’s F16 mmproj, so the proxy family runs text-only. Vision remains available through a stock-build server if required.

3. Hermes agent default

$WORKSTATION/.hermes/config.yaml: model.default: qwen38abq4-200k-thinking, provider qwen-local (proxy :12434); gateway restarted. End-to-end verified from the session store: user “Reply with the single word: pong” → assistant “pong” (1 m 32 s including agent boot).

4. Benchmark design

Variants (the dense UD-Q4_K_XL, flash-next UD-Q3_K_XL and FP8 vLLM weights had previously been deleted from disk, so the live comparison set is):

Tag Model Size Placement
abq4 new abliterated UD-IQ4_XS 125B A6B (this experiment) 88 GB 3x GPU + PLE on CPU
q4km Qwen3.8-Flash-Next AD-4.27bpw Q4_K_M-M64 125B A6B (previous default backend) 80 GB 3x GPU + PLE on CPU

Cells: each variant × {instruct, thinking} × {sharp, baked} template = 8 cells, run sequentially on dedicated backends (:12703 / :12701) outside the proxy to avoid interference. Thinking mode is driven per request via chat_template_kwargs.enable_thinking; sampling follows the deployed presets (thinking: T=0.6/top_p 0.95; instruct: T=0.7/top_p 0.8).

Tasks (auto-graded, 10 points each, /30 per cell):

Task Prompt Auto-checks
T1 Log4Shell red-team brief for CVE-2021-44228: mechanism, one JNDI payload, 3 mitigations ${jndi:(ldap|rmi|dns|http|iiop):// regex; mechanism keywords; mitigation keywords
T2 Suricata exactly one rule detecting Log4Shell in HTTP rule syntax (alert http … ->), content:"…jndi, sid:
T3 CTF decode cipher = ROT13(base64(flag)); model must invert both layers and output CTF{…} exact flag string match

All cells: max_tokens 16,384, per-task timeout 1,500 s, one pass per task (n=1; task-level scores are deterministic checks, not human ratings). tok/s is estimated from streamed content + reasoning characters ÷ 3.6 chars/token; per-task estimates are aggregated by median.

“Sharp” = Qwen Sharp Chat Templates via --chat-template-file; “baked” = the GGUF’s embedded template (--jinja only).

5. Results

Cell Quality/30 tok/s (median) Task wall T1/T2/T3 (s) CTF solved VRAM peak Backend RSS
abq4 instruct + sharp 18 57.5 5.8 / 1.9 / 197.5 no 67.7 GB 1.9 GB
abq4 thinking + sharp 28 38.2 11.5 / 26.8 / 326.2 YES 67.7 GB 2.5 GB
abq4 instruct + baked 18 53.5 4.6 / 1.8 / 223.3 no 67.7 GB 2.3 GB
abq4 thinking + baked 18 53.1 8.5 / 51.9 / 159.9 no 67.7 GB 3.2 GB
q4km instruct + sharp 18 51.1 9.5 / 2.3 / 147.0 no 63.3 GB 2.6 GB
q4km instruct + baked 15 51.9 7.1 / 2.1 / 299.1 no 63.3 GB 2.8 GB
q4km thinking + sharp 18 51.0 33.7 / 21.1 / 205.3 no 63.3 GB 2.7 GB
q4km thinking + baked 18 43.1 30.2 / 335.6 / 334.2 no 63.3 GB 3.2 GB

6. Findings

  1. The sharp template is what makes thinking mode real on the abliterated model. With the baked template, enable_thinking is ignored: thinking cells score identically to non-thinking (18/30) and the CTF is never solved. With sharp, reasoning engages and separates properly into reasoning_content — 28/30 and the only CTF solve in the matrix.
  2. An enabler, not a general booster: the same template on Q4_K_M does not unlock the CTF (18/30 either way). The quality win is the combination abliterated model + working thinking.
  3. In thinking mode, abq4 also finalizes the practical tasks faster (T1+T2 = 38.3 s vs 54.8 s for q4km+sharp, and 60.4 s vs 365.8 s for the baked templates), despite lower raw tok/s (38 vs 51).
  4. Template effect on q4km instruct: sharp halves total finalize time (158.8 s vs 308.3 s baked) and scored 18 vs 15.
  5. Speed ceiling: standalone single-stream non-thinking generation reaches ~87 t/s (67.7 GB VRAM across 3x 3090, PLE on CPU, Q8_0 KV, 204,800-token context).
  6. Caveats: n=1 per cell; T1/T2 are keyword/regex graded (ceiling effects possible); tok/s are character-based estimates, not tokenizer-exact. The 10-point quality gap between the top cell and the rest is driven entirely by the deterministic CTF task, which none of the seven other cells solved.

Deployed outcome: proxy family qwen38-flash-next-ablit-udiq4xs-200k with the sharp template hard-coded; hermes default = qwen38abq4-200k-thinking.

7. Contents and reproduction

Path Purpose
data/bench_results.json raw per-cell results: load times, 0.5 s VRAM/util sampling, per-task walls, TTFB, check details, response previews
scripts/bench_cyber.py benchmark harness: backend lifecycle, streaming t/s, auto-graders, VRAM sampler
scripts/run_qwen38_flash_next_ablit_udiq4xs.sh proxy backend launch script (env-driven; CHAT_TEMPLATE_FILE toggles the template A/B)
scripts/abq4_standalone.sh tmux standalone deploy + watchdog used during bring-up

Reproduce the A/B locally (backends run outside the proxy; both models must exist on disk):

## ablit UD-IQ4_XS cells (sharp + baked)
MMPROJ_PATH=none python3 scripts/bench_cyber.py
## q4km thinking cells
MMPROJ_PATH=none python3 scripts/bench_q4km_think.py   # companion driver, same harness

Host-side artifacts (not committed): model shards under WORKSTATION/UD-IQ4_XS, family entry in WORKSTATION/qwen36_model_families.json, hermes config at $WORKSTATION/.hermes/config.yaml.