Back to the experimentSupporting notebook

Qwen3.8-Flash-Next UD-Q4_K_XL: CPU-MoE offload deployment + speed tuning (3x3090, 200k ctx)

Source: qwen-next-flash/udq4kxl-cpu-moe-200k-2026-09-20/README.md · revision 6550ead3945b

1. Download

aria2c hard-caps -x at 16 connections/server (165 impossible). Parallelism via 3 tmux sessions, one part each: aria2c -x16 -s16 -k8M -c <url>. All 4 parts verified byte-exact vs HF metadata (part1 10,946,624 B … part4 12,087,983,520 B).

Model path points at part1; llama.cpp concatenates multipart GGUFs transparently.

2. The load problem: 104 GiB vs 72 GB VRAM

Obligatory CPU offload of ~25 GiB of MoE experts. Key findings:

  1. -ncmoe N + -sm layer is a trap: layer-split ignores weight; with first-N experts on CPU, GPU0 (early, near-empty blocks) idles while GPU1/2 get overloaded contiguous buffers (repeatable failed to allocate CUDA1 buffer of size 15788813824).
  2. -ot separator is COMMA, not semicolon — semicolons give unknown buffer type.
  3. LOAD_MODE=none is REQUIRED with CPU -ot rules: mmap + CPU overrides double-allocate.
  4. Historical tensor names from run_qwen38_flash_next_q4km.sh.bak-20260905 still valid: blk.N.ffn_{gate|up|down}_exps.weight, per_layer_token_embd.weight (PLE).
  5. Explicit per-GPU expert-block ranges are the only reliable placement; everything else rides the script’s baseline (-ngl 999 -sm layer -ts 1,1,1 -fa on --repack --op-offload -b 16384 -ub 400 -t 14 --cpu-strict -Cr 0-13).

3. Speed A/B: expert blocks on GPU (204800 ctx, q8_0 KV, 300-token streaming bench)

GPU expert blocks VRAM GiB (0/1/2) CPU blocks gen tok/s
8/8/8 (24 total) 17.3 / 16.0 / 16.6 24 18.9
10/10/10 (30) 20.3 / 19.0 / 19.6 18 22.1
11/11/11 (33) 21.8 / 20.5 / 21.3 15 26.1
12/12/12 (36) — winner 23.3 / 22.0 / 22.8 12 31.1
12/13/12 (37) 23.3 / 23.5 / 22.8 11 30.9 (plateau, tighter GPU1)

Balanced splits beat unbalanced at equal block count (a 12/3/7 accidental split gave only 19.5 — pipeline stalls on the empty GPU dominate). 12/12/12 = same speed as 12/13/12 with better margins → production choice.

4. Verdict & context

5. Files

Not yet wired into qwen36_model_families.json (proxy family pending decision).