Journal

Pipeline parallelism, INT4, and a 262K context

The patched vLLM stack measured 67.4 tokens per second for one stream and 73.1 aggregate for two.

Vlad / experimentos.
The boundary

Vision and MTP were disabled. The llama.cpp comparison used a different quant and context configuration.

The question

Pipeline parallelism and INT4 offered a different serving route for the large model. This record qualifies its long context, single-stream decode and two-stream aggregate throughput, with vision and MTP disabled.

From the original notebook

  • Date: 2026-09-20
  • Host: 3x RTX 3090 (72 GB VRAM), 125 GB RAM, Ubuntu 24.04, ext4 NVMe
  • Model: Minachist/Qwen3.8-Flash-Next-INT4-Mixed-AutoRound @ rev 1703ff595285561b141648609291377d3faf3b64 (164 GiB, compressed-tensors WNA16/Marlin)
  • Stack: podman 4.9.3 + nvidia-container-toolkit (CDI), vLLM 0.29.1rc1.dev47+gdc36fcce9 pinned by image digest sha256:43f13b4c…47bb96, upstream vllm-patch/3x3090 (vllm patch + PLE host-gather + model-state-hook → FULL_DECODE_ONLY CUDA graphs)
  • Serving: flash-next container on localhost, behind qwen36-multi-proxy :12434 as family flashnext-int4-vllm-262k (backend_managed=false)
  • Result: 67.4 tok/s 1-stream, 73.1 tok/s 2-stream aggregate — 2.2x the best llama.cpp UD-Q4_K_XL config (31.1 tok/s)

Config (upstream compose replicated verbatim; all load-bearing)

PP=3 / TP=1, --language-model-only, --engram-config {cpu_offload:true}, max-model-len 262144, max-num-seqs 4, max-num-batched-tokens 512, gpu-mem-util 0.97, kv-cache-memory 550000000 (do NOT round up — allocator power-of-2 boundary; ~560 MB → 63→111 GiB host RAM), prefix caching OFF, cudagraph_mode=FULL_DECODE_ONLY sizes [1,2,4]. Env: PLE mmap, BREAKABLE_CUDAGRAPH=0, QSA_KV_OFFLOAD=1 (MAX_GIB 56, ARENA 100663296), expandable_segments, ESTIMATE_CUDAGRAPHS=0.

Validation gates — all exact upstream numbers

Gate Measured
GPU KV cache 1,067,300 tokens (4.07x of 262,144)
QSA block size 12144 (not 1568 → patch applied)
QSA GPU footprint 66 B/token/layer
Host pool 171 blocks x 3.961 GiB pinned, 12 QSA layers
PLE mmap 96 GiB on NVMe (WORKSTATION/ngram_table.bin)
VRAM 23.4 GiB x 3
Host RAM ~46 GiB still available under load

Pitfalls hit (the real findings)

  1. podman-compose 1.0.6 silently drops CDI devices: [nvidia.com/gpu=all] — container starts, vLLM dies with Failed to infer device type. podman inspect .HostConfig.Devices was []. Fix: plain podman run (scripts/run-flash-next.sh) instead of compose.
  2. PLE bind-mount needs a pre-existing file at the exact size — podman creates a directory when the host path is missing, and the patched loader rejects a 0-byte file as stale. Fix: truncate -s 102400491520 WORKSTATION/ngram_table.bin (sparse) before first launch; loader fills it.
  3. GPU squatters OOM rank 0 at 0.97 util — a respawning ComfyUI (256 MiB) broke PP0 while PP1/2 loaded. Weights load 21.55 GiB/rank, only ~2 GiB headroom. Kill/stop everything on the GPUs first.
  4. Patched image build is fast (~3 min); the 164 GiB download ~30 min (hf_xet ~100 MB/s), first boot ~10 min.

Ops

  • Start: WORKSTATION/run-flash-next.sh; logs: sudo podman logs -f flash-next
  • Stop: sudo podman rm -f flash-next
  • Bench: scripts/bench_vllm.py [n_streams]
  • Proxy: family flashnext-int4-vllm-262k on :12434 (backend_managed=false → proxy will NOT kill/restart it)

Context

Replaced (same day): UD-Q4_K_XL GGUF (104 GiB, best 31.1 tok/s — see udq4kxl-cpu-moe-200k-2026-09-20), exl3-3.75bpw and the BF16 source checkpoint were deleted to make room. Old GGUF wrapper preserved at WORKSTATION/qwen38-gguf-backup. MTP/vision deliberately OFF per upstream (3-GPU build omits the 4.86 GiB MTP module); n-gram drafting unsupported on this execution path.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS