Journal

From a model download to a working local service

The deployment notebook records the EXL3 serving family, cache choices and real runtime constraints.

Vlad / experimentos.
The boundary

This notebook contains historical configurations; use the dated experiments to interpret each measurement.

The question

A downloaded model does not yet make a useful service. This deployment notebook follows formats, cache settings, routing and resource constraints; its dated experiments explain which configuration produced each result.

From the original notebook

Also see the in-progress official-BF16 EXL3 campaign: exl3-350hq-375-2026-09-15 (3.50 HQ + 3.75 MTP6, plan + journey).

udq4kxl-cpu-moe-200k-2026-09-20 — UD-Q4_K_XL GGUF (104 GiB) CPU-MoE deployment: -ot expert placement tuning, 12/12/12 GPU split, 31.1 tok/s @ 204800/q8.

vllm-int4-autoround-3x3090-2026-09-20 — INT4-Mixed-AutoRound via patched vLLM PP3 podman deploy: 67.4 tok/s 1-stream / 73 aggregate @ 262144. New champion.

Qwen3.8-Flash-Next abliterated UD-IQ4_XS: quantization, deployment and template A/B

  • Date: 2026-09-06 (single-day run)
  • Host: laguna-s21 — 3x RTX 3090 (72 GB aggregate VRAM), 128 GB RAM, Ubuntu Linux
  • Inference stack: llama.cpp flash-next fork at WORKSTATION/llama.cpp-flash-next (--repack, --op-offload, -Cr CPU pinning), plus the stock ggml-org build for the quantization pipeline
  • Serving: qwen36-multi-proxy on :12434, backends on loopback
  • Agent: hermes, default provider qwen-local

This experiment covers the complete pipeline for moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 → Unsloth-style UD-IQ4_XS GGUF: torchao-FP8 conversion, live agentic importance-matrix collection, proxy family deployment as the hermes agent default, and an 8-cell benchmark (2 model variants × thinking on/off × chat template sharp/baked) on auto-graded cybersecurity tasks at a 16,384-token completion cap.

1. Quantization job (quant-jobs/fp8-unc-udiq4xs-6h)

Step Detail
Source moltriver/Qwen3.8-Flash-Next-Uncensored-FP8 (abliterated 125B A6B MoE, native block-FP8 via torchao, 512 experts/layer, PLE n-gram table 320M rows × 160 dims)
Converter patches torchao layout handling: per-expert sharded .weight_scale and per-layer scales for the PLE table (mandatory, otherwise scale lookup fails)
Master q8_0 re-quantization of all 1224 tensors → 188.2 GB single-file master
Importance matrix 170/200 chunks of live agentic traffic (2 h timebox), collected with CPU-MoE expert split (-ts 1,1,1; 16 GPU layer-blocks / 32 CPU layer-blocks, PLE on CPU), merged with an AtomicChat-calibration imatrix (4,000 chunks) → 579 MB file, 99.8% expert coverage at collection end
Quantization llama-quantize with the Unsloth UD-IQ4_XS tensor-type map (643 overrides; experts mostly iq3_s, PLE table iq4_nl, attention Q/K/MMLP at higher precision)
Output 88 GB, 41 shards → WORKSTATION/Qwen38FN-Unc-FP8-unsloth-UD-IQ4_XS-00001-of-00041.gguf
Smoke test (stock build, CPU experts) prompt 75.6 t/s, generation 21.8 t/s, correct reply

2. Proxy deployment (clone of the UD-Q4 family pattern)

Serving parameters were cloned from the qwen38-unsloth-q4 (UD-Q4_K_XL) family: sharp chat-template hook, Q8_0 KV cache, 204,800-token context, batch 16,384 / ubatch 400, 14 pinned CPU threads (0-13), metrics enabled. Hardware placement follows the existing flash-next families because the model (88 GB) exceeds VRAM (72 GB):

  • PLE table on CPU — per_layer_token_embd (29 GB of IQ4_NL) pinned via -ot '^per_layer_token_embd\.weight$=CPU'; all other tensors GPU-resident across 3 cards.
  • The flash-next fork with --repack --op-offload is ~4x faster than the stock-build CPU-expert layout used during the quant job (86.7 vs 21.8 t/s single-stream generation).
  • Family qwen38-flash-next-ablit-udiq4xs-200k, backend localhost, variants instruct / general / coding / thinking exposed as qwen38abq4-200k-* on :12434.

Root cause found during bring-up

Requests through the proxy hung indefinitely (clients saw HTTP 000/timeouts) while the backend itself was healthy. _backend_is_ready() in the proxy requires /v1/models to advertise exactly the family’s backend_model string. Our launch script used --alias qwen38abq4-200k while the family key is qwen38-flash-next-ablit-udiq4xs-200k → never “ready” → the GPU-acquisition wait (900 s) deadlocked on VRAM held by the proxy’s own backend. Fix: SERVED_MODEL_NAME must equal backend_model (the convention every other family follows). After the fix, a cold proxy launch from warm page cache answers its first request in ~17 s.

Standalone verification (before moving to the proxy)

A tmux-deployed standalone server on 0.0.0.0:12703 with a watchdog (hard deadline 1,500 s, 10 s progress trail, automatic self-test) was used to validate behaviour in isolation: ready in 21 s from warm page cache, self-test reply correct, 67.7 GB VRAM, 86.7 t/s single-stream non-thinking generation. See scripts/abq4_standalone.sh.

Known limitation: vision

The flash-next fork’s mtmd loader aborts (ggml_backend_buffer_set_usage in clip_model_loader::load_tensors) when loading this model’s F16 mmproj, so the proxy family runs text-only. Vision remains available through a stock-build server if required.

3. Hermes agent default

$WORKSTATION/.hermes/config.yaml: model.default: qwen38abq4-200k-thinking, provider qwen-local (proxy :12434); gateway restarted. End-to-end verified from the session store: user “Reply with the single word: pong” → assistant “pong” (1 m 32 s including agent boot).

4. Benchmark design

Variants (the dense UD-Q4_K_XL, flash-next UD-Q3_K_XL and FP8 vLLM weights had previously been deleted from disk, so the live comparison set is):

Tag Model Size Placement
abq4 new abliterated UD-IQ4_XS 125B A6B (this experiment) 88 GB 3x GPU + PLE on CPU
q4km Qwen3.8-Flash-Next AD-4.27bpw Q4_K_M-M64 125B A6B (previous default backend) 80 GB 3x GPU + PLE on CPU

Cells: each variant × {instruct, thinking} × {sharp, baked} template = 8 cells, run sequentially on dedicated backends (:12703 / :12701) outside the proxy to avoid interference. Thinking mode is driven per request via chat_template_kwargs.enable_thinking; sampling follows the deployed presets (thinking: T=0.6/top_p 0.95; instruct: T=0.7/top_p 0.8).

Tasks (auto-graded, 10 points each, /30 per cell):

Task Prompt Auto-checks
T1 Log4Shell red-team brief for CVE-2021-44228: mechanism, one JNDI payload, 3 mitigations ${jndi:(ldap|rmi|dns|http|iiop):// regex; mechanism keywords; mitigation keywords
T2 Suricata exactly one rule detecting Log4Shell in HTTP rule syntax (alert http … ->), content:"…jndi, sid:
T3 CTF decode cipher = ROT13(base64(flag)); model must invert both layers and output CTF{…} exact flag string match

All cells: max_tokens 16,384, per-task timeout 1,500 s, one pass per task (n=1; task-level scores are deterministic checks, not human ratings). tok/s is estimated from streamed content + reasoning characters ÷ 3.6 chars/token; per-task estimates are aggregated by median.

“Sharp” = Qwen Sharp Chat Templates via --chat-template-file; “baked” = the GGUF’s embedded template (--jinja only).

5. Results

Cell Quality/30 tok/s (median) Task wall T1/T2/T3 (s) CTF solved VRAM peak Backend RSS
abq4 instruct + sharp 18 57.5 5.8 / 1.9 / 197.5 no 67.7 GB 1.9 GB
abq4 thinking + sharp 28 38.2 11.5 / 26.8 / 326.2 YES 67.7 GB 2.5 GB
abq4 instruct + baked 18 53.5 4.6 / 1.8 / 223.3 no 67.7 GB 2.3 GB
abq4 thinking + baked 18 53.1 8.5 / 51.9 / 159.9 no 67.7 GB 3.2 GB
q4km instruct + sharp 18 51.1 9.5 / 2.3 / 147.0 no 63.3 GB 2.6 GB
q4km instruct + baked 15 51.9 7.1 / 2.1 / 299.1 no 63.3 GB 2.8 GB
q4km thinking + sharp 18 51.0 33.7 / 21.1 / 205.3 no 63.3 GB 2.7 GB
q4km thinking + baked 18 43.1 30.2 / 335.6 / 334.2 no 63.3 GB 3.2 GB

6. Findings

  1. The sharp template is what makes thinking mode real on the abliterated model. With the baked template, enable_thinking is ignored: thinking cells score identically to non-thinking (18/30) and the CTF is never solved. With sharp, reasoning engages and separates properly into reasoning_content — 28/30 and the only CTF solve in the matrix.
  2. An enabler, not a general booster: the same template on Q4_K_M does not unlock the CTF (18/30 either way). The quality win is the combination abliterated model + working thinking.
  3. In thinking mode, abq4 also finalizes the practical tasks faster (T1+T2 = 38.3 s vs 54.8 s for q4km+sharp, and 60.4 s vs 365.8 s for the baked templates), despite lower raw tok/s (38 vs 51).
  4. Template effect on q4km instruct: sharp halves total finalize time (158.8 s vs 308.3 s baked) and scored 18 vs 15.
  5. Speed ceiling: standalone single-stream non-thinking generation reaches ~87 t/s (67.7 GB VRAM across 3x 3090, PLE on CPU, Q8_0 KV, 204,800-token context).
  6. Caveats: n=1 per cell; T1/T2 are keyword/regex graded (ceiling effects possible); tok/s are character-based estimates, not tokenizer-exact. The 10-point quality gap between the top cell and the rest is driven entirely by the deterministic CTF task, which none of the seven other cells solved.

Deployed outcome: proxy family qwen38-flash-next-ablit-udiq4xs-200k with the sharp template hard-coded; hermes default = qwen38abq4-200k-thinking.

7. Contents and reproduction

Path Purpose
data/bench_results.json raw per-cell results: load times, 0.5 s VRAM/util sampling, per-task walls, TTFB, check details, response previews
scripts/bench_cyber.py benchmark harness: backend lifecycle, streaming t/s, auto-graders, VRAM sampler
scripts/run_qwen38_flash_next_ablit_udiq4xs.sh proxy backend launch script (env-driven; CHAT_TEMPLATE_FILE toggles the template A/B)
scripts/abq4_standalone.sh tmux standalone deploy + watchdog used during bring-up

Reproduce the A/B locally (backends run outside the proxy; both models must exist on disk):

## ablit UD-IQ4_XS cells (sharp + baked)
MMPROJ_PATH=none python3 scripts/bench_cyber.py
## q4km thinking cells
MMPROJ_PATH=none python3 scripts/bench_q4km_think.py   # companion driver, same harness

Host-side artifacts (not committed): model shards under WORKSTATION/UD-IQ4_XS, family entry in WORKSTATION/qwen36_model_families.json, hermes config at $WORKSTATION/.hermes/config.yaml.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS