Journal

Making room for a 200K context

The warm proxy decode median was 59.17 tokens per second; a 198,960-token request passed with GPU headroom.

Vlad / experimentos.
The boundary

The qualified one-slot profile is specific to three RTX 3090s and its exact quantization and cache configuration.

The question

Getting a large model to load is the first hurdle; completing a nearly full context is another. This experiment qualified a specific three-GPU profile with a warm decode measurement and a real long-context request.

From the original notebook

This experiment establishes a stable 204,800-token llama.cpp serving profile for unsloth/Qwen3.8-Flash-Next-GGUF UD-Q3_K_XL on three 24 GB RTX 3090s. The endpoint was exercised through an OpenAI-compatible lifecycle proxy, not only against a directly started server.

Retained profile

  • llama.cpp CUDA build commit 511f9c1;
  • one server slot, layer split, and tensor split 1,1,1;
  • all MoE expert tensors on CUDA, with per_layer_token_embd.weight on CPU;
  • layer 4 expert placement: gate/up on CUDA1 and down on CUDA2;
  • Q8_0 K and V cache, flash attention, 204,800-token context;
  • logical batch 16,384, microbatch 400;
  • 14 generation and batch threads restricted by llama.cpp to logical CPUs 0-13, representing seven physical P-cores with SMT;
  • mmap loading, disabled lazy tensor reads, tensor repacking, and operation offload.

The exact launcher is archived at scripts/run_qwen38_flash_next_udq3xl.sh.

Measured results

Workload Result Boundary
Warm proxy decode, median 59.17 tok/s Six 256-token requests, three prose and three code
Sustained proxy decode 57.51 tok/s One 512-token request
Warm 18K prompt prefill 1,095.55 tok/s Two uncached requests through the proxy
198,960-token prompt prefill 469.86 tok/s One uncached proxy-path qualification request
Full-context minimum free VRAM 1,826 / 1,578 / 1,584 MiB GPU0 / GPU1 / GPU2, 500 ms sampling
Swap during qualification si=0, so=0 One-second vmstat samples

The 198,960-token request returned exactly QUALIFICATION_OK; the proxy reported 198,960 prompt tokens and four completion tokens. This demonstrates that the retained one-slot profile fits and completes the stated context on this workstation. It is not evidence of concurrent 200K-request capacity, accuracy, or performance on other GPUs or builds.

The raw, machine-readable summary is data/qualification.json. Its values are taken from the exact proxy response and concurrent nvidia-smi and vmstat monitors. The monitors are intentionally not committed because they contain machine-local timestamp traces rather than reusable experimental inputs.

Selection process

The prior CPU-MoE recipe decoded at 42.7 tok/s. Moving all routed experts to the GPUs and placing only the PLE table on CPU improved the retained warm decode result to 59.17 tok/s, a 38.6% increase calculated as (59.17 / 42.7 - 1) * 100.

Several near-neighbor settings were rejected:

Candidate Outcome
Tensor split mode 32.88 tok/s decode; rejected
Row split mode Unsupported by the CUDA split-buffer implementation
--no-op-offload No decode improvement; rejected
--load-mode none More resident RAM and lower real-server decode; rejected
--no-repack Lower decode and prefill; rejected
Microbatch 448 1,446 / 1,452 MiB free at 198,960 tokens on GPUs 1/2; failed the 1.5 GiB floor

The microbatch was reduced to 400. That final change passed the full-context memory floor while retaining 59.17 tok/s median warm decode.

Reproduction

The launcher has workstation-specific defaults for the llama.cpp build and the three-part GGUF model. Override those paths and then start it normally:

bash scripts/run_qwen38_flash_next_udq3xl.sh

Use one slot and an uncached request when validating the 200K boundary. Sample all GPUs at 500 ms or faster, verify the minimum free VRAM on every card, and monitor vmstat for swap activity. The retained safety floor is 1,536 MiB of free VRAM per GPU.

Limits

  • Results are from one i7-12700KF and three RTX 3090s with the listed build.
  • The full-context run has one qualification sample, not a confidence interval.
  • Prefill and decode figures use different prompt lengths and are not directly interchangeable.
  • The profile has one slot; it does not establish multi-user throughput.
  • The model’s GGUF metadata identified an A3B, 176.94B-parameter model during benchmarking; this report does not infer architecture details beyond that observed metadata.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (1)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS