Back to the experimentSupporting notebook

Qwen3.8-Flash-Next INT4-Mixed-AutoRound via vLLM PP3 (3x3090, 262k ctx)

Source: qwen-next-flash/vllm-int4-autoround-3x3090-2026-09-20/README.md · revision 6550ead3945b

Config (upstream compose replicated verbatim; all load-bearing)

PP=3 / TP=1, --language-model-only, --engram-config {cpu_offload:true}, max-model-len 262144, max-num-seqs 4, max-num-batched-tokens 512, gpu-mem-util 0.97, kv-cache-memory 550000000 (do NOT round up — allocator power-of-2 boundary; ~560 MB → 63→111 GiB host RAM), prefix caching OFF, cudagraph_mode=FULL_DECODE_ONLY sizes [1,2,4]. Env: PLE mmap, BREAKABLE_CUDAGRAPH=0, QSA_KV_OFFLOAD=1 (MAX_GIB 56, ARENA 100663296), expandable_segments, ESTIMATE_CUDAGRAPHS=0.

Validation gates — all exact upstream numbers

Gate Measured
GPU KV cache 1,067,300 tokens (4.07x of 262,144)
QSA block size 12144 (not 1568 → patch applied)
QSA GPU footprint 66 B/token/layer
Host pool 171 blocks x 3.961 GiB pinned, 12 QSA layers
PLE mmap 96 GiB on NVMe (WORKSTATION/ngram_table.bin)
VRAM 23.4 GiB x 3
Host RAM ~46 GiB still available under load

Pitfalls hit (the real findings)

  1. podman-compose 1.0.6 silently drops CDI devices: [nvidia.com/gpu=all] — container starts, vLLM dies with Failed to infer device type. podman inspect .HostConfig.Devices was []. Fix: plain podman run (scripts/run-flash-next.sh) instead of compose.
  2. PLE bind-mount needs a pre-existing file at the exact size — podman creates a directory when the host path is missing, and the patched loader rejects a 0-byte file as stale. Fix: truncate -s 102400491520 WORKSTATION/ngram_table.bin (sparse) before first launch; loader fills it.
  3. GPU squatters OOM rank 0 at 0.97 util — a respawning ComfyUI (256 MiB) broke PP0 while PP1/2 loaded. Weights load 21.55 GiB/rank, only ~2 GiB headroom. Kill/stop everything on the GPUs first.
  4. Patched image build is fast (~3 min); the 164 GiB download ~30 min (hf_xet ~100 MB/s), first boot ~10 min.

Ops

Context

Replaced (same day): UD-Q4_K_XL GGUF (104 GiB, best 31.1 tok/s — see udq4kxl-cpu-moe-200k-2026-09-20), exl3-3.75bpw and the BF16 source checkpoint were deleted to make room. Old GGUF wrapper preserved at WORKSTATION/qwen38-gguf-backup. MTP/vision deliberately OFF per upstream (3-GPU build omits the 4.86 GiB MTP module); n-gram drafting unsupported on this execution path.