The question
Getting a large model to load is the first hurdle; completing a nearly full context is another. This experiment qualified a specific three-GPU profile with a warm decode measurement and a real long-context request.
From the original notebook
This experiment establishes a stable 204,800-token llama.cpp serving profile
for unsloth/Qwen3.8-Flash-Next-GGUF UD-Q3_K_XL on three 24 GB RTX 3090s.
The endpoint was exercised through an OpenAI-compatible lifecycle proxy, not
only against a directly started server.
Retained profile
- llama.cpp CUDA build commit
511f9c1; - one server slot, layer split, and tensor split
1,1,1; - all MoE expert tensors on CUDA, with
per_layer_token_embd.weighton CPU; - layer 4 expert placement: gate/up on CUDA1 and down on CUDA2;
- Q8_0 K and V cache, flash attention, 204,800-token context;
- logical batch 16,384, microbatch 400;
- 14 generation and batch threads restricted by llama.cpp to logical CPUs
0-13, representing seven physical P-cores with SMT; - mmap loading, disabled lazy tensor reads, tensor repacking, and operation offload.
The exact launcher is archived at
scripts/run_qwen38_flash_next_udq3xl.sh.
Measured results
| Workload | Result | Boundary |
|---|---|---|
| Warm proxy decode, median | 59.17 tok/s | Six 256-token requests, three prose and three code |
| Sustained proxy decode | 57.51 tok/s | One 512-token request |
| Warm 18K prompt prefill | 1,095.55 tok/s | Two uncached requests through the proxy |
| 198,960-token prompt prefill | 469.86 tok/s | One uncached proxy-path qualification request |
| Full-context minimum free VRAM | 1,826 / 1,578 / 1,584 MiB | GPU0 / GPU1 / GPU2, 500 ms sampling |
| Swap during qualification | si=0, so=0 | One-second vmstat samples |
The 198,960-token request returned exactly QUALIFICATION_OK; the proxy
reported 198,960 prompt tokens and four completion tokens. This demonstrates
that the retained one-slot profile fits and completes the stated context on
this workstation. It is not evidence of concurrent 200K-request capacity,
accuracy, or performance on other GPUs or builds.
The raw, machine-readable summary is
data/qualification.json. Its values are taken
from the exact proxy response and concurrent nvidia-smi and vmstat
monitors. The monitors are intentionally not committed because they contain
machine-local timestamp traces rather than reusable experimental inputs.
Selection process
The prior CPU-MoE recipe decoded at 42.7 tok/s. Moving all routed experts to
the GPUs and placing only the PLE table on CPU improved the retained warm
decode result to 59.17 tok/s, a 38.6% increase calculated as
(59.17 / 42.7 - 1) * 100.
Several near-neighbor settings were rejected:
| Candidate | Outcome |
|---|---|
| Tensor split mode | 32.88 tok/s decode; rejected |
| Row split mode | Unsupported by the CUDA split-buffer implementation |
--no-op-offload |
No decode improvement; rejected |
--load-mode none |
More resident RAM and lower real-server decode; rejected |
--no-repack |
Lower decode and prefill; rejected |
| Microbatch 448 | 1,446 / 1,452 MiB free at 198,960 tokens on GPUs 1/2; failed the 1.5 GiB floor |
The microbatch was reduced to 400. That final change passed the full-context memory floor while retaining 59.17 tok/s median warm decode.
Reproduction
The launcher has workstation-specific defaults for the llama.cpp build and the three-part GGUF model. Override those paths and then start it normally:
bash scripts/run_qwen38_flash_next_udq3xl.sh
Use one slot and an uncached request when validating the 200K boundary. Sample
all GPUs at 500 ms or faster, verify the minimum free VRAM on every card, and
monitor vmstat for swap activity. The retained safety floor is 1,536 MiB of
free VRAM per GPU.
Limits
- Results are from one i7-12700KF and three RTX 3090s with the listed build.
- The full-context run has one qualification sample, not a confidence interval.
- Prefill and decode figures use different prompt lengths and are not directly interchangeable.
- The profile has one slot; it does not establish multi-user throughput.
- The model’s GGUF metadata identified an A3B, 176.94B-parameter model during benchmarking; this report does not infer architecture details beyond that observed metadata.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.