The question
Pipeline parallelism and INT4 offered a different serving route for the large model. This record qualifies its long context, single-stream decode and two-stream aggregate throughput, with vision and MTP disabled.
From the original notebook
- Date: 2026-09-20
- Host: 3x RTX 3090 (72 GB VRAM), 125 GB RAM, Ubuntu 24.04, ext4 NVMe
- Model:
Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound@ rev1703ff595285561b141648609291377d3faf3b64(164 GiB, compressed-tensors WNA16/Marlin) - Stack: podman 4.9.3 + nvidia-container-toolkit (CDI), vLLM
0.29.1rc1.dev47+gdc36fcce9pinned by image digestsha256:43f13b4c…47bb96, upstream vllm-patch/3x3090 (vllm patch + PLE host-gather + model-state-hook → FULL_DECODE_ONLY CUDA graphs) - Serving:
flash-nextcontainer onlocalhost, behindqwen36-multi-proxy :12434as familyflashnext-int4-vllm-262k(backend_managed=false) - Result: 67.4 tok/s 1-stream, 73.1 tok/s 2-stream aggregate — 2.2x the best llama.cpp UD-Q4_K_XL config (31.1 tok/s)
Config (upstream compose replicated verbatim; all load-bearing)
PP=3 / TP=1, --language-model-only, --engram-config {cpu_offload:true}, max-model-len 262144,
max-num-seqs 4, max-num-batched-tokens 512, gpu-mem-util 0.97, kv-cache-memory 550000000 (do NOT round up —
allocator power-of-2 boundary; ~560 MB → 63→111 GiB host RAM), prefix caching OFF,
cudagraph_mode=FULL_DECODE_ONLY sizes [1,2,4]. Env: PLE mmap, BREAKABLE_CUDAGRAPH=0, QSA_KV_OFFLOAD=1
(MAX_GIB 56, ARENA 100663296), expandable_segments, ESTIMATE_CUDAGRAPHS=0.
Validation gates — all exact upstream numbers
| Gate | Measured |
|---|---|
| GPU KV cache | 1,067,300 tokens (4.07x of 262,144) |
| QSA block size | 12144 (not 1568 → patch applied) |
| QSA GPU footprint | 66 B/token/layer |
| Host pool | 171 blocks x 3.961 GiB pinned, 12 QSA layers |
| PLE mmap | 96 GiB on NVMe (WORKSTATION/ngram_table.bin) |
| VRAM | 23.4 GiB x 3 |
| Host RAM | ~46 GiB still available under load |
Pitfalls hit (the real findings)
- podman-compose 1.0.6 silently drops CDI
devices: [nvidia.com/gpu=all]— container starts, vLLM dies withFailed to infer device type.podman inspect .HostConfig.Deviceswas[]. Fix: plainpodman run(scripts/run-flash-next.sh) instead of compose. - PLE bind-mount needs a pre-existing file at the exact size — podman creates a directory when the host
path is missing, and the patched loader rejects a 0-byte file as stale. Fix:
truncate -s 102400491520 WORKSTATION/ngram_table.bin(sparse) before first launch; loader fills it. - GPU squatters OOM rank 0 at 0.97 util — a respawning ComfyUI (256 MiB) broke PP0 while PP1/2 loaded. Weights load 21.55 GiB/rank, only ~2 GiB headroom. Kill/stop everything on the GPUs first.
- Patched image build is fast (~3 min); the 164 GiB download ~30 min (hf_xet ~100 MB/s), first boot ~10 min.
Ops
- Start:
WORKSTATION/run-flash-next.sh; logs:sudo podman logs -f flash-next - Stop:
sudo podman rm -f flash-next - Bench:
scripts/bench_vllm.py [n_streams] - Proxy: family
flashnext-int4-vllm-262kon:12434(backend_managed=false → proxy will NOT kill/restart it)
Context
Replaced (same day): UD-Q4_K_XL GGUF (104 GiB, best 31.1 tok/s — see
udq4kxl-cpu-moe-200k-2026-09-20), exl3-3.75bpw and the BF16 source
checkpoint were deleted to make room. Old GGUF wrapper preserved at WORKSTATION/qwen38-gguf-backup.
MTP/vision deliberately OFF per upstream (3-GPU build omits the 4.86 GiB MTP module); n-gram drafting
unsupported on this execution path.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.