- Date: 2026-09-20 (single-day run)
- Host: 3x RTX 3090 (72 GB VRAM), 125 GB RAM, i7 CPU
- Model:
unsloth/Qwen3.8-Flash-Next-GGUFUD-Q4_K_XL — 4-part multipart GGUF, 104 GiB total (~130B MoE, 48 blocks, 512 experts/10 active, only 12 full-attention blocks) - Inference: llama.cpp flash-next fork
WORKSTATION/llama-server(build 0.3.0-dev, commit 9723942), launched viaWORKSTATION/run_qwen38_flash_next_q4km.sh - Serving: port 12705, name
flashnext-udq4kxl-200k, ctx 204800, KV q8_0 - Winner: expert blocks 12/12/12 per GPU + 12 blocks CPU + PLE on CPU → 31.1 tok/s
1. Download
aria2c hard-caps -x at 16 connections/server (165 impossible). Parallelism via 3 tmux sessions, one part each:
aria2c -x16 -s16 -k8M -c <url>. All 4 parts verified byte-exact vs HF metadata (part1 10,946,624 B … part4 12,087,983,520 B).
Model path points at part1; llama.cpp concatenates multipart GGUFs transparently.
2. The load problem: 104 GiB vs 72 GB VRAM
Obligatory CPU offload of ~25 GiB of MoE experts. Key findings:
-ncmoe N+-sm layeris a trap: layer-split ignores weight; with first-N experts on CPU, GPU0 (early, near-empty blocks) idles while GPU1/2 get overloaded contiguous buffers (repeatablefailed to allocate CUDA1 buffer of size 15788813824).-otseparator is COMMA, not semicolon — semicolons giveunknown buffer type.LOAD_MODE=noneis REQUIRED with CPU-otrules: mmap + CPU overrides double-allocate.- Historical tensor names from
run_qwen38_flash_next_q4km.sh.bak-20260905still valid:blk.N.ffn_{gate|up|down}_exps.weight,per_layer_token_embd.weight(PLE). - Explicit per-GPU expert-block ranges are the only reliable placement; everything else rides
the script’s baseline (
-ngl 999 -sm layer -ts 1,1,1 -fa on --repack --op-offload -b 16384 -ub 400 -t 14 --cpu-strict -Cr 0-13).
3. Speed A/B: expert blocks on GPU (204800 ctx, q8_0 KV, 300-token streaming bench)
| GPU expert blocks | VRAM GiB (0/1/2) | CPU blocks | gen tok/s |
|---|---|---|---|
| 8/8/8 (24 total) | 17.3 / 16.0 / 16.6 | 24 | 18.9 |
| 10/10/10 (30) | 20.3 / 19.0 / 19.6 | 18 | 22.1 |
| 11/11/11 (33) | 21.8 / 20.5 / 21.3 | 15 | 26.1 |
| 12/12/12 (36) — winner | 23.3 / 22.0 / 22.8 | 12 | 31.1 |
| 12/13/12 (37) | 23.3 / 23.5 / 22.8 | 11 | 30.9 (plateau, tighter GPU1) |
Balanced splits beat unbalanced at equal block count (a 12/3/7 accidental split gave only 19.5 — pipeline stalls on the empty GPU dominate). 12/12/12 = same speed as 12/13/12 with better margins → production choice.
4. Verdict & context
- 31 tok/s is the physics ceiling for this quant on this host: ~25 GiB of experts must run on CPU
(
--op-offload). Reference points: old all-GPU Q4_K_M M64 did 60.5 tok/s @ 204800/q8; exl3-4bpw (all-GPU, 65.1 GB weights) did 67.6 tok/s @ 265984/q8, 66.3 @ 307200/q6. - exl3-4bpw beats UD-Q4_K_XL ~2x on speed AND reaches larger ctx, but costs 102 GB disk vs 104 GB and lacks GGUF ecosystem flexibility. This GGUF run is the llama.cpp-native deployment.
- Model fine at 204800/q8 with ~0.7 GiB margin on GPU0; higher ctx would need KV q6 or fewer GPU blocks.
5. Files
scripts/run_udq4xl.sh— winning launch wrapper (12/12/12 + PLE→CPU + LOAD_MODE=none)scripts/bench_gen.py— streaming gen tok/s benchmark- Live server: tmux/setsid, log
/tmp/udq4xl-server.log, healthcurl localhost/health
Not yet wired into qwen36_model_families.json (proxy family pending decision).