Back to the experimentSupporting notebook

Qwen3.8-Flash-Next UD-Q3_K_XL on three RTX 3090s

Source: qwen38/flash-next-udq3xl-3x3090-200k-2026-08-29/README.md · revision 6550ead3945b

This experiment establishes a stable 204,800-token llama.cpp serving profile for unsloth/Qwen3.8-Flash-Next-GGUF UD-Q3_K_XL on three 24 GB RTX 3090s. The endpoint was exercised through an OpenAI-compatible lifecycle proxy, not only against a directly started server.

Retained profile

The exact launcher is archived at scripts/run_qwen38_flash_next_udq3xl.sh.

Measured results

Workload Result Boundary
Warm proxy decode, median 59.17 tok/s Six 256-token requests, three prose and three code
Sustained proxy decode 57.51 tok/s One 512-token request
Warm 18K prompt prefill 1,095.55 tok/s Two uncached requests through the proxy
198,960-token prompt prefill 469.86 tok/s One uncached proxy-path qualification request
Full-context minimum free VRAM 1,826 / 1,578 / 1,584 MiB GPU0 / GPU1 / GPU2, 500 ms sampling
Swap during qualification si=0, so=0 One-second vmstat samples

The 198,960-token request returned exactly QUALIFICATION_OK; the proxy reported 198,960 prompt tokens and four completion tokens. This demonstrates that the retained one-slot profile fits and completes the stated context on this workstation. It is not evidence of concurrent 200K-request capacity, accuracy, or performance on other GPUs or builds.

The raw, machine-readable summary is data/qualification.json. Its values are taken from the exact proxy response and concurrent nvidia-smi and vmstat monitors. The monitors are intentionally not committed because they contain machine-local timestamp traces rather than reusable experimental inputs.

Selection process

The prior CPU-MoE recipe decoded at 42.7 tok/s. Moving all routed experts to the GPUs and placing only the PLE table on CPU improved the retained warm decode result to 59.17 tok/s, a 38.6% increase calculated as (59.17 / 42.7 - 1) * 100.

Several near-neighbor settings were rejected:

Candidate Outcome
Tensor split mode 32.88 tok/s decode; rejected
Row split mode Unsupported by the CUDA split-buffer implementation
--no-op-offload No decode improvement; rejected
--load-mode none More resident RAM and lower real-server decode; rejected
--no-repack Lower decode and prefill; rejected
Microbatch 448 1,446 / 1,452 MiB free at 198,960 tokens on GPUs 1/2; failed the 1.5 GiB floor

The microbatch was reduced to 400. That final change passed the full-context memory floor while retaining 59.17 tok/s median warm decode.

Reproduction

The launcher has workstation-specific defaults for the llama.cpp build and the three-part GGUF model. Override those paths and then start it normally:

bash scripts/run_qwen38_flash_next_udq3xl.sh

Use one slot and an uncached request when validating the 200K boundary. Sample all GPUs at 500 ms or faster, verify the minimum free VRAM on every card, and monitor vmstat for swap activity. The retained safety floor is 1,536 MiB of free VRAM per GPU.

Limits