The question
A downloaded model does not yet make a useful service. This deployment notebook follows formats, cache settings, routing and resource constraints; its dated experiments explain which configuration produced each result.
From the original notebook
Also see the in-progress official-BF16 EXL3 campaign: exl3-350hq-375-2026-09-15 (3.50 HQ + 3.75 MTP6, plan + journey).
udq4kxl-cpu-moe-200k-2026-09-20 — UD-Q4_K_XL GGUF (104 GiB) CPU-MoE deployment: -ot expert placement tuning, 12/12/12 GPU split, 31.1 tok/s @ 204800/q8.
vllm-int4-autoround-3x3090-2026-09-20 — INT4-Mixed-AutoRound via patched vLLM PP3 podman deploy: 67.4 tok/s 1-stream / 73 aggregate @ 262144. New champion.
Qwen3.8-Flash-Next abliterated UD-IQ4_XS: quantization, deployment and template A/B
- Date: 2026-09-06 (single-day run)
- Host:
laguna-s21— 3x RTX 3090 (72 GB aggregate VRAM), 128 GB RAM, Ubuntu Linux - Inference stack: llama.cpp flash-next fork at
WORKSTATION/llama.cpp-flash-next(--repack,--op-offload,-CrCPU pinning), plus the stock ggml-org build for the quantization pipeline - Serving:
qwen36-multi-proxyon:12434, backends on loopback - Agent: hermes, default provider
qwen-local
This experiment covers the complete pipeline for
moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 → Unsloth-style UD-IQ4_XS GGUF:
torchao-FP8 conversion, live agentic importance-matrix collection, proxy family deployment as
the hermes agent default, and an 8-cell benchmark (2 model variants × thinking on/off × chat
template sharp/baked) on auto-graded cybersecurity tasks at a 16,384-token completion cap.
1. Quantization job (quant-jobs/fp8-unc-udiq4xs-6h)
| Step | Detail |
|---|---|
| Source | moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 (abliterated 125B A6B MoE, native block-FP8 via torchao, 512 experts/layer, PLE n-gram table 320M rows × 160 dims) |
| Converter patches | torchao layout handling: per-expert sharded .weight_scale and per-layer scales for the PLE table (mandatory, otherwise scale lookup fails) |
| Master | q8_0 re-quantization of all 1224 tensors → 188.2 GB single-file master |
| Importance matrix | 170/200 chunks of live agentic traffic (2 h timebox), collected with CPU-MoE expert split (-ts 1,1,1; 16 GPU layer-blocks / 32 CPU layer-blocks, PLE on CPU), merged with an AtomicChat-calibration imatrix (4,000 chunks) → 579 MB file, 99.8% expert coverage at collection end |
| Quantization | llama-quantize with the Unsloth UD-IQ4_XS tensor-type map (643 overrides; experts mostly iq3_s, PLE table iq4_nl, attention Q/K/MMLP at higher precision) |
| Output | 88 GB, 41 shards → WORKSTATION/Qwen38FN-Unc-FP8-unsloth-UD-IQ4_XS-00001-of-00041.gguf |
| Smoke test (stock build, CPU experts) | prompt 75.6 t/s, generation 21.8 t/s, correct reply |
2. Proxy deployment (clone of the UD-Q4 family pattern)
Serving parameters were cloned from the qwen38-unsloth-q4 (UD-Q4_K_XL) family: sharp
chat-template hook, Q8_0 KV cache, 204,800-token context, batch 16,384 / ubatch 400, 14 pinned
CPU threads (0-13), metrics enabled. Hardware placement follows the existing flash-next
families because the model (88 GB) exceeds VRAM (72 GB):
- PLE table on CPU —
per_layer_token_embd(29 GB of IQ4_NL) pinned via-ot '^per_layer_token_embd\.weight$=CPU'; all other tensors GPU-resident across 3 cards. - The flash-next fork with
--repack --op-offloadis ~4x faster than the stock-build CPU-expert layout used during the quant job (86.7 vs 21.8 t/s single-stream generation). - Family
qwen38-flash-next-ablit-udiq4xs-200k, backendlocalhost, variantsinstruct / general / coding / thinkingexposed asqwen38abq4-200k-*on:12434.
Root cause found during bring-up
Requests through the proxy hung indefinitely (clients saw HTTP 000/timeouts) while the backend
itself was healthy. _backend_is_ready() in the proxy requires /v1/models to advertise
exactly the family’s backend_model string. Our launch script used
--alias qwen38abq4-200k while the family key is
qwen38-flash-next-ablit-udiq4xs-200k → never “ready” → the GPU-acquisition wait (900 s)
deadlocked on VRAM held by the proxy’s own backend. Fix: SERVED_MODEL_NAME must equal
backend_model (the convention every other family follows). After the fix, a cold proxy
launch from warm page cache answers its first request in ~17 s.
Standalone verification (before moving to the proxy)
A tmux-deployed standalone server on 0.0.0.0:12703 with a watchdog (hard deadline 1,500 s,
10 s progress trail, automatic self-test) was used to validate behaviour in isolation:
ready in 21 s from warm page cache, self-test reply correct, 67.7 GB VRAM, 86.7 t/s
single-stream non-thinking generation. See scripts/abq4_standalone.sh.
Known limitation: vision
The flash-next fork’s mtmd loader aborts (ggml_backend_buffer_set_usage in
clip_model_loader::load_tensors) when loading this model’s F16 mmproj, so the proxy family
runs text-only. Vision remains available through a stock-build server if required.
3. Hermes agent default
$WORKSTATION/.hermes/config.yaml: model.default: qwen38abq4-200k-thinking, provider qwen-local
(proxy :12434); gateway restarted. End-to-end verified from the session store:
user “Reply with the single word: pong” → assistant “pong” (1 m 32 s including agent boot).
4. Benchmark design
Variants (the dense UD-Q4_K_XL, flash-next UD-Q3_K_XL and FP8 vLLM weights had previously been deleted from disk, so the live comparison set is):
| Tag | Model | Size | Placement |
|---|---|---|---|
abq4 |
new abliterated UD-IQ4_XS 125B A6B (this experiment) | 88 GB | 3x GPU + PLE on CPU |
q4km |
Qwen3.8-Flash-Next AD-4.27bpw Q4_K_M-M64 125B A6B (previous default backend) | 80 GB | 3x GPU + PLE on CPU |
Cells: each variant × {instruct, thinking} × {sharp, baked} template = 8 cells, run
sequentially on dedicated backends (:12703 / :12701) outside the proxy to avoid
interference. Thinking mode is driven per request via chat_template_kwargs.enable_thinking;
sampling follows the deployed presets (thinking: T=0.6/top_p 0.95; instruct: T=0.7/top_p 0.8).
Tasks (auto-graded, 10 points each, /30 per cell):
| Task | Prompt | Auto-checks |
|---|---|---|
| T1 Log4Shell | red-team brief for CVE-2021-44228: mechanism, one JNDI payload, 3 mitigations | ${jndi:(ldap|rmi|dns|http|iiop):// regex; mechanism keywords; mitigation keywords |
| T2 Suricata | exactly one rule detecting Log4Shell in HTTP | rule syntax (alert http … ->), content:"…jndi, sid: |
| T3 CTF decode | cipher = ROT13(base64(flag)); model must invert both layers and output CTF{…} |
exact flag string match |
All cells: max_tokens 16,384, per-task timeout 1,500 s, one pass per task (n=1; task-level
scores are deterministic checks, not human ratings). tok/s is estimated from streamed content +
reasoning characters ÷ 3.6 chars/token; per-task estimates are aggregated by median.
“Sharp” = Qwen Sharp Chat Templates
via --chat-template-file; “baked” = the GGUF’s embedded template (--jinja only).
5. Results
| Cell | Quality/30 | tok/s (median) | Task wall T1/T2/T3 (s) | CTF solved | VRAM peak | Backend RSS |
|---|---|---|---|---|---|---|
| abq4 instruct + sharp | 18 | 57.5 | 5.8 / 1.9 / 197.5 | no | 67.7 GB | 1.9 GB |
| abq4 thinking + sharp | 28 | 38.2 | 11.5 / 26.8 / 326.2 | YES | 67.7 GB | 2.5 GB |
| abq4 instruct + baked | 18 | 53.5 | 4.6 / 1.8 / 223.3 | no | 67.7 GB | 2.3 GB |
| abq4 thinking + baked | 18 | 53.1 | 8.5 / 51.9 / 159.9 | no | 67.7 GB | 3.2 GB |
| q4km instruct + sharp | 18 | 51.1 | 9.5 / 2.3 / 147.0 | no | 63.3 GB | 2.6 GB |
| q4km instruct + baked | 15 | 51.9 | 7.1 / 2.1 / 299.1 | no | 63.3 GB | 2.8 GB |
| q4km thinking + sharp | 18 | 51.0 | 33.7 / 21.1 / 205.3 | no | 63.3 GB | 2.7 GB |
| q4km thinking + baked | 18 | 43.1 | 30.2 / 335.6 / 334.2 | no | 63.3 GB | 3.2 GB |
6. Findings
- The sharp template is what makes thinking mode real on the abliterated model. With the
baked template,
enable_thinkingis ignored: thinking cells score identically to non-thinking (18/30) and the CTF is never solved. With sharp, reasoning engages and separates properly intoreasoning_content— 28/30 and the only CTF solve in the matrix. - An enabler, not a general booster: the same template on Q4_K_M does not unlock the CTF (18/30 either way). The quality win is the combination abliterated model + working thinking.
- In thinking mode, abq4 also finalizes the practical tasks faster (T1+T2 = 38.3 s vs 54.8 s for q4km+sharp, and 60.4 s vs 365.8 s for the baked templates), despite lower raw tok/s (38 vs 51).
- Template effect on q4km instruct: sharp halves total finalize time (158.8 s vs 308.3 s baked) and scored 18 vs 15.
- Speed ceiling: standalone single-stream non-thinking generation reaches ~87 t/s (67.7 GB VRAM across 3x 3090, PLE on CPU, Q8_0 KV, 204,800-token context).
- Caveats: n=1 per cell; T1/T2 are keyword/regex graded (ceiling effects possible); tok/s are character-based estimates, not tokenizer-exact. The 10-point quality gap between the top cell and the rest is driven entirely by the deterministic CTF task, which none of the seven other cells solved.
Deployed outcome: proxy family qwen38-flash-next-ablit-udiq4xs-200k with the sharp
template hard-coded; hermes default = qwen38abq4-200k-thinking.
7. Contents and reproduction
| Path | Purpose |
|---|---|
data/bench_results.json |
raw per-cell results: load times, 0.5 s VRAM/util sampling, per-task walls, TTFB, check details, response previews |
scripts/bench_cyber.py |
benchmark harness: backend lifecycle, streaming t/s, auto-graders, VRAM sampler |
scripts/run_qwen38_flash_next_ablit_udiq4xs.sh |
proxy backend launch script (env-driven; CHAT_TEMPLATE_FILE toggles the template A/B) |
scripts/abq4_standalone.sh |
tmux standalone deploy + watchdog used during bring-up |
Reproduce the A/B locally (backends run outside the proxy; both models must exist on disk):
## ablit UD-IQ4_XS cells (sharp + baked)
MMPROJ_PATH=none python3 scripts/bench_cyber.py
## q4km thinking cells
MMPROJ_PATH=none python3 scripts/bench_q4km_think.py # companion driver, same harness
Host-side artifacts (not committed): model shards under
WORKSTATION/UD-IQ4_XS, family entry in
WORKSTATION/qwen36_model_families.json, hermes config at
$WORKSTATION/.hermes/config.yaml.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.