Journal

DFlash2 on a single RTX 3090

Ornith measured 210.16 tokens per second on the realistic 256-token prompt and 270.40 on predictable prose.

Vlad / experimentos.
The boundary

Prompt predictability changes draft acceptance; these are different workloads, not interchangeable throughput estimates.

The question

Speculative decoding needs to earn its speedup on more than an easy completion. This single-GPU DFlash2 record separates predictable prose from a realistic prompt and records the effect of draft acceptance.

From the original notebook

Purpose

This report records the exact build, deployment parameters, validation path, and performance measurements used to deploy Ornith 1.5 35B A3B Abliterated with a BF16 DFlash2 draft model on one RTX 3090.

The experiment used the DFlash2 implementation from llama.cpp pull request 27342. The tested source base is commit 2f3923bc81346046aa5765dda15fb28497f49ac6. Older installed llama.cpp builds rejected the draft checkpoint with wrong number of tensors; expected 96, got 69 because they only understood the earlier DFlash tensor layout.

Artifacts

Target model:

WORKSTATION/Ornith-1.5-35B-A3B-Abliterated-Q4_K_M.gguf

Target SHA256:

a07f299e83a398b5078c1cb8ab4ec96333c8ea9d10d0cb479cb073056fedd3d0

DFlash2 draft model:

WORKSTATION/Ornith-1.5-35B-A3B-jzinno_DFlash2-BF16.gguf

Draft SHA256:

e2713e376ac5e5f1ec06a9d0a57351bb603a3130a84e385c072ffd8dbe307af0

The draft loader reported:

spec_type=draft-dflash
block_size=16
mask_token_id=248077
n_extract=8
sample_from_anchor=true

Build

scripts/ornith/build-dflash2-cuda.sh

The tested build used CUDA 12.0, compute capability 8.6, Release mode, and LLAMA_CURL=OFF. The generated server identified as llama.cpp build 10513, commit 2f3923bc8.

Host environment:

Kernel:        Linux 6.8.0-137-generic x86_64
GPU:           3 x NVIDIA GeForce RTX 3090
NVIDIA driver: 595.58.03
CMake:         3.28.3
CUDA compiler: 12.0
G++:           13.3.0

Production parameters

GPU count:                 1
GPU index:                 0
Target context:            160000
Parallel slots:            1
Target GPU layers:         99
Flash attention:           on
Target K cache:            q8_0
Target V cache:            q8_0
Batch size:                1024
Micro-batch size:          128
CPU threads:               1
HTTP threads:              1
Fit:                       off
Speculative type:          draft-dflash
Draft GPU layers:          all
Draft maximum tokens:      7
Draft minimum tokens:      0
Draft probability minimum: 0
Multimodal projector:      disabled

Start the standalone server with:

GPU_INDEX=0 scripts/ornith/run-ornith-dflash2.sh

The 160K profile is close to the 24 GB memory ceiling. The production worker used about 24070 MiB on GPU 0 and left about 58 MiB free after loading and generation. The profile completed real 256-token requests, but it has almost no extra allocation margin.

Vision is intentionally disabled. The experimental DFlash2 branch did not provide a reliable multimodal path, and the projector did not fit safely beside the target, draft, and 160K Q8 KV cache.

Proxy integration

The public proxy kept the existing aliases:

ornith-instruct
ornith-general
ornith-coding
ornith-thinking

The local proxy integration was committed separately in WORKSTATION/NAMEOFMODEL:

43f1cbc deploy Ornith DFlash2 speculative backend
eaa37c5 preserve public model prefixes for variant families

The family points to the DFlash2-capable server binary, the target GGUF, and the BF16 draft GGUF. It advertises 160000 context tokens, one GPU, Q8 K/V cache, and text-only input. Run scripts/ornith/check-proxy.sh to verify the user-visible route.

Benchmarks

All decode samples used one RTX 3090, 160K configured context for DFlash2 tests, Q8 K/V, one slot, draft maximum 7, temperature 0, prompt-cache disabled for each request, and forced 256 or 512 generated tokens.

The raw measurements are stored in ornith-dflash2-2026-08-27.csv. Accepted and drafted token counters came from each request’s server timing record. Cumulative metrics were checked against the per-request totals.

Realistic Python prompt, 256 tokens

Prompt:

Write a Python function that finds duplicate integers in a list and explain its complexity.
GPU Power limit Decode samples, tok/s Median Draft acceptance
GPU 0 220 W 208.98, 210.16, 210.64 210.16 64.20%
GPU 2 270 W 212.82, 219.17, 218.76 218.76 60.77%

GPU 2 was 4.09% faster by median despite lower draft acceptance.

Predictable counting prompt, 512 tokens

GPU Decode samples, tok/s Median Draft acceptance
GPU 0 270.43, 270.35, 270.40 270.40 87.28%
GPU 2 258.12, 257.77, 260.40 258.12 74.65%

GPU 2 was 4.54% slower because the generated path produced lower speculative acceptance. DFlash2 acceptance had more effect than the additional 50 W power allowance.

Non-speculative control

At 8K context with speculative decoding disabled, GPU 2 produced 155.95, 156.47, and 154.01 tok/s for the Python prompt. The median was 155.95 tok/s. A prior identical GPU 0 sample was 143.53 tok/s, suggesting roughly 8.7% raw decode benefit from GPU 2, but the GPU 0 control had only one recorded sample and is not a robust median comparison.

GPU 2 reached 257-269 W and approximately 1.82 GHz during the benchmark. GPU 0 was power-limited to 220 W.

Findings

  1. The exact BF16 DFlash2 checkpoint works with the DFlash2 pull-request implementation.
  2. Draft maximum 7 was more stable than 15. Draft maximum 15 produced very high best cases but severe slowdowns on low-acceptance prose.
  3. Realistic DFlash2 speed depends strongly on acceptance. More GPU power did not guarantee higher end-to-end throughput.
  4. The 160K Q8 KV profile fits one 24 GB RTX 3090 but leaves minimal memory headroom.
  5. GPU 0 remained the production choice because GPU 2 is shared by the two-GPU Hauhau service and its realistic DFlash2 gain was only about 4%.
  6. The Hauhau service was stopped only after its active-request count reached zero and was restored after benchmarking.

Validation performed

  • Exact draft download SHA256 matched the remote artifact hash.
  • DFlash2 target and draft loaded together on GPU 0 and GPU 2.
  • Standalone /health, /v1/models, /v1/completions, and speculative metrics were checked.
  • Public proxy :12434 returned the exact text DFLASH2 LIVE through ornith-instruct.
  • A 256-token public request completed at 189.34 tok/s with 67.31% draft acceptance in that sample.
  • All four existing Ornith aliases advertised 160000 context tokens.
  • The production service and the temporarily displaced Hauhau service were healthy with zero active requests after testing.

Reproduce

scripts/ornith/build-dflash2-cuda.sh
GPU_INDEX=2 PORT=12822 scripts/ornith/run-ornith-dflash2.sh
BASE_URL=private workstation service MODEL=ornith-dflash2 RUNS=3 MAX_TOKENS=256 scripts/ornith/benchmark-dflash2.sh
tests/test-ornith-dflash2-scripts.sh
BASE_URL=private workstation service MODEL=ornith-dflash2 tests/test-ornith-dflash2-runtime.sh

Do not run the 160K profile on a GPU that has another allocation. Check process ownership and active requests before stopping a managed workload.

Keep exploring

Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.

Follow the evidence

From notebook to finding.

This story is based on the archived experiment at revision 6550ead3945b. Original timestamps, workloads and qualification limits belong to that record.

Original GitHub record
Supporting notebooks (2)

GitHub source links require access to the private archive. The readable notes and aggregate chart exports are included here.

Back to the journal Follow via RSS