The question
Speculative decoding needs to earn its speedup on more than an easy completion. This single-GPU DFlash2 record separates predictable prose from a realistic prompt and records the effect of draft acceptance.
From the original notebook
Purpose
This report records the exact build, deployment parameters, validation path, and performance measurements used to deploy Ornith 1.5 35B A3B Abliterated with a BF16 DFlash2 draft model on one RTX 3090.
The experiment used the DFlash2 implementation from llama.cpp pull request 27342. The tested source base is commit 2f3923bc81346046aa5765dda15fb28497f49ac6. Older installed llama.cpp builds rejected the draft checkpoint with wrong number of tensors; expected 96, got 69 because they only understood the earlier DFlash tensor layout.
Artifacts
Target model:
WORKSTATION/Ornith-1.5-35B-A3B-Abliterated-Q4_K_M.gguf
Target SHA256:
a07f299e83a398b5078c1cb8ab4ec96333c8ea9d10d0cb479cb073056fedd3d0
DFlash2 draft model:
WORKSTATION/Ornith-1.5-35B-A3B-jzinno_DFlash2-BF16.gguf
Draft SHA256:
e2713e376ac5e5f1ec06a9d0a57351bb603a3130a84e385c072ffd8dbe307af0
The draft loader reported:
spec_type=draft-dflash
block_size=16
mask_token_id=248077
n_extract=8
sample_from_anchor=true
Build
scripts/ornith/build-dflash2-cuda.sh
The tested build used CUDA 12.0, compute capability 8.6, Release mode, and LLAMA_CURL=OFF. The generated server identified as llama.cpp build 10513, commit 2f3923bc8.
Host environment:
Kernel: Linux 6.8.0-137-generic x86_64
GPU: 3 x NVIDIA GeForce RTX 3090
NVIDIA driver: 595.58.03
CMake: 3.28.3
CUDA compiler: 12.0
G++: 13.3.0
Production parameters
GPU count: 1
GPU index: 0
Target context: 160000
Parallel slots: 1
Target GPU layers: 99
Flash attention: on
Target K cache: q8_0
Target V cache: q8_0
Batch size: 1024
Micro-batch size: 128
CPU threads: 1
HTTP threads: 1
Fit: off
Speculative type: draft-dflash
Draft GPU layers: all
Draft maximum tokens: 7
Draft minimum tokens: 0
Draft probability minimum: 0
Multimodal projector: disabled
Start the standalone server with:
GPU_INDEX=0 scripts/ornith/run-ornith-dflash2.sh
The 160K profile is close to the 24 GB memory ceiling. The production worker used about 24070 MiB on GPU 0 and left about 58 MiB free after loading and generation. The profile completed real 256-token requests, but it has almost no extra allocation margin.
Vision is intentionally disabled. The experimental DFlash2 branch did not provide a reliable multimodal path, and the projector did not fit safely beside the target, draft, and 160K Q8 KV cache.
Proxy integration
The public proxy kept the existing aliases:
ornith-instruct
ornith-general
ornith-coding
ornith-thinking
The local proxy integration was committed separately in WORKSTATION/NAMEOFMODEL:
43f1cbc deploy Ornith DFlash2 speculative backend
eaa37c5 preserve public model prefixes for variant families
The family points to the DFlash2-capable server binary, the target GGUF, and the BF16 draft GGUF. It advertises 160000 context tokens, one GPU, Q8 K/V cache, and text-only input. Run scripts/ornith/check-proxy.sh to verify the user-visible route.
Benchmarks
All decode samples used one RTX 3090, 160K configured context for DFlash2 tests, Q8 K/V, one slot, draft maximum 7, temperature 0, prompt-cache disabled for each request, and forced 256 or 512 generated tokens.
The raw measurements are stored in ornith-dflash2-2026-08-27.csv. Accepted and drafted token counters came from each request’s server timing record. Cumulative metrics were checked against the per-request totals.
Realistic Python prompt, 256 tokens
Prompt:
Write a Python function that finds duplicate integers in a list and explain its complexity.
| GPU | Power limit | Decode samples, tok/s | Median | Draft acceptance |
|---|---|---|---|---|
| GPU 0 | 220 W | 208.98, 210.16, 210.64 | 210.16 | 64.20% |
| GPU 2 | 270 W | 212.82, 219.17, 218.76 | 218.76 | 60.77% |
GPU 2 was 4.09% faster by median despite lower draft acceptance.
Predictable counting prompt, 512 tokens
| GPU | Decode samples, tok/s | Median | Draft acceptance |
|---|---|---|---|
| GPU 0 | 270.43, 270.35, 270.40 | 270.40 | 87.28% |
| GPU 2 | 258.12, 257.77, 260.40 | 258.12 | 74.65% |
GPU 2 was 4.54% slower because the generated path produced lower speculative acceptance. DFlash2 acceptance had more effect than the additional 50 W power allowance.
Non-speculative control
At 8K context with speculative decoding disabled, GPU 2 produced 155.95, 156.47, and 154.01 tok/s for the Python prompt. The median was 155.95 tok/s. A prior identical GPU 0 sample was 143.53 tok/s, suggesting roughly 8.7% raw decode benefit from GPU 2, but the GPU 0 control had only one recorded sample and is not a robust median comparison.
GPU 2 reached 257-269 W and approximately 1.82 GHz during the benchmark. GPU 0 was power-limited to 220 W.
Findings
- The exact BF16 DFlash2 checkpoint works with the DFlash2 pull-request implementation.
- Draft maximum 7 was more stable than 15. Draft maximum 15 produced very high best cases but severe slowdowns on low-acceptance prose.
- Realistic DFlash2 speed depends strongly on acceptance. More GPU power did not guarantee higher end-to-end throughput.
- The 160K Q8 KV profile fits one 24 GB RTX 3090 but leaves minimal memory headroom.
- GPU 0 remained the production choice because GPU 2 is shared by the two-GPU Hauhau service and its realistic DFlash2 gain was only about 4%.
- The Hauhau service was stopped only after its active-request count reached zero and was restored after benchmarking.
Validation performed
- Exact draft download SHA256 matched the remote artifact hash.
- DFlash2 target and draft loaded together on GPU 0 and GPU 2.
- Standalone
/health,/v1/models,/v1/completions, and speculative metrics were checked. - Public proxy
:12434returned the exact textDFLASH2 LIVEthroughornith-instruct. - A 256-token public request completed at 189.34 tok/s with 67.31% draft acceptance in that sample.
- All four existing Ornith aliases advertised 160000 context tokens.
- The production service and the temporarily displaced Hauhau service were healthy with zero active requests after testing.
Reproduce
scripts/ornith/build-dflash2-cuda.sh
GPU_INDEX=2 PORT=12822 scripts/ornith/run-ornith-dflash2.sh
BASE_URL=private workstation service MODEL=ornith-dflash2 RUNS=3 MAX_TOKENS=256 scripts/ornith/benchmark-dflash2.sh
tests/test-ornith-dflash2-scripts.sh
BASE_URL=private workstation service MODEL=ornith-dflash2 tests/test-ornith-dflash2-runtime.sh
Do not run the 160K profile on a GPU that has another allocation. Check process ownership and active requests before stopping a managed workload.
Keep exploring
Original experiment record. Workstation paths have been generalized. Detailed measurements below retain their original workload and validation boundaries.