This experiment compares the stock Qwen3.8 chat templates with
Qwen Sharp Chat Templates v22.4.0
at medium and xhigh reasoning effort. It covers three live two-GPU deployments:
- HauhauCS Aggressive
Q6_K_PGGUF with llama.cpp FastMTP; - OrcaRouter uncensored FP8 with vLLM MTP;
- official Unsloth
Qwen3.8-27B-UD-Q6_K_XLGGUF with llama.cpp FastMTP.
The run matrix contains 12 complete 100-item runs. Every raw response, reasoning trace, token
count, finish reason, paired comparison, suite item, and scoring script is included under
data/ and scripts/.
Read the publication-ready blog post or open the self-contained interactive chart report. The updated report includes accuracy, wall time, completion tokens, prompt-plus-completion totals, average response length, median per-response tok/s, aggregate tok/s at concurrency 2, and matched Sharp-versus-stock savings.
Deployment outcome
Correctness was the first selection criterion and observed completion time was the tiebreaker. The retained proxy configuration is:
| Deployment | Retained template | Recommended effort | Why |
|---|---|---|---|
| Hauhau Q6 | Sharp v22.4.0 | xhigh when maximum benchmark correctness matters; medium for latency |
Sharp scored 180/200 across both efforts versus 179/200 stock and used 24.7% less combined wall time. Its best single run was 91/100 at xhigh. |
| Unsloth UD-Q6_K_XL | Sharp v22.4.0 | medium |
Both Sharp efforts scored 90/100, but medium was 2.36x faster and used about half the completion tokens. Sharp scored 180/200 versus 178/200 stock. |
| OrcaRouter FP8 | Sharp v22.4.0 operational default | medium |
Sharp medium is the deployed speed/token-efficiency choice. Stock xhigh reached 90/100 versus 88/100 for Sharp, but the two-answer difference was not statistically significant. |
These one- and two-answer differences are operational choices, not proof of a general quality winner: no paired accuracy contrast was statistically significant in this 100-item sample.
Complete results
All runs used concurrency 2 and a 16,000-token maximum completion allowance. Wall time is end-to-end for the complete 100-item run and can include a cold model load or first-shape JIT.
| Model | Template | Effort | Correct | Wall seconds | Completion tokens | Median response tok/s | Aggregate tok/s (c=2) | Unparsed | Length stops |
|---|---|---|---|---|---|---|---|---|---|
| Hauhau Q6 | Sharp | medium | 89/100 | 918.211 | 64,036 | 34.21 | 69.74 | 0 | 0 |
| Hauhau Q6 | Sharp | xhigh | 91/100 | 2,546.913 | 170,550 | 33.77 | 66.96 | 2 | 2 |
| Hauhau Q6 | Stock | medium | 90/100 | 1,206.307 | 89,209 | 36.30 | 73.95 | 0 | 0 |
| Hauhau Q6 | Stock | xhigh | 89/100 | 3,395.125 | 206,147 | 30.79 | 60.72 | 4 | 5 |
| Unsloth Q6 | Sharp | medium | 90/100 | 913.819 | 65,189 | 34.79 | 71.34 | 0 | 0 |
| Unsloth Q6 | Sharp | xhigh | 90/100 | 2,155.354 | 130,852 | 30.22 | 60.71 | 1 | 1 |
| Unsloth Q6 | Stock | medium | 90/100 | 1,032.878 | 76,516 | 37.17 | 74.08 | 0 | 0 |
| Unsloth Q6 | Stock | xhigh | 88/100 | 2,763.000 | 168,773 | 30.57 | 61.08 | 3 | 4 |
| FP8 | Sharp | medium | 88/100 | 481.396 | 59,255 | 61.66 | 123.09 | 0 | 0 |
| FP8 | Sharp | xhigh | 88/100 | 1,541.898 | 149,974 | 52.56 | 97.27 | 4 | 4 |
| FP8 | Stock | medium | 88/100 | 616.227 | 82,979 | 68.26 | 134.66 | 0 | 0 |
| FP8 | Stock | xhigh | 90/100 | 1,679.790 | 166,235 | 54.59 | 98.96 | 4 | 4 |
All 12 runs had zero transport/backend errors. A response that reached 16K without a parseable
FINAL: answer remained wrong; no manual credit was added.
What xhigh changed
Using the retained template for each model:
| Model/template | Accuracy change, medium to xhigh | Wall-time multiplier | Completion-token multiplier |
|---|---|---|---|
| Hauhau / Sharp | +2 answers | 2.77x | 2.66x |
| Unsloth / Sharp | 0 answers | 2.36x | 2.01x |
| FP8 / Stock | +2 answers | 2.73x | 2.00x |
The observed +2 changes for Hauhau and FP8 were not significant: paired bootstrap intervals included zero and exact McNemar p-values were 0.727 and 0.688, respectively. Unsloth Sharp was an accuracy tie with McNemar p=1.0.
Template A/B paired correctness
The table reports Sharp minus stock on the same 100 item IDs.
| Model | Effort | Stock | Sharp | Stock-only correct | Sharp-only correct | Difference | Exact McNemar p |
|---|---|---|---|---|---|---|---|
| Hauhau Q6 | medium | 90 | 89 | 1 | 0 | -1 point | 1.000 |
| Hauhau Q6 | xhigh | 89 | 91 | 1 | 3 | +2 points | 0.625 |
| Unsloth Q6 | medium | 90 | 90 | 2 | 2 | 0 points | 1.000 |
| Unsloth Q6 | xhigh | 88 | 90 | 1 | 3 | +2 points | 0.625 |
| FP8 | medium | 88 | 88 | 2 | 2 | 0 points | 1.000 |
| FP8 | xhigh | 90 | 88 | 2 | 0 | -2 points | 0.500 |
Correctness protocol
The frozen suite has 100 examples:
- 40 MMLU-Pro multiple-choice questions;
- 20 ARC-Challenge questions;
- 20 HellaSwag questions;
- 20 GSM8K numeric questions.
Prompts, source row IDs, gold answers, seeds, and source provenance are stored in
data/suite.json. Multiple-choice responses had to end with
FINAL: <LETTER> and GSM8K responses with FINAL: <NUMBER>. The scorer normalizes basic numeric
formatting and uses the final matching answer. It does not judge free-form reasoning or rescue a
truncated answer manually.
Both template families explicitly support reasoning_effort=medium|xhigh; the effort value was
passed through chat_template_kwargs with thinking enabled and preserved. The public thinking
routes forced the same effective sampling preset in every arm: temperature 1.0, top-p 0.95,
top-k 20, presence penalty 1.5, and min-p 0.0. The runner used two worker threads and fixed seeds.
Runtime profiles and the new Unsloth deployment
The exact retained family definitions are in
configs/final-model-profiles.json.
The newly added public route family is:
qwen38-uq6-instructqwen38-uq6-generalqwen38-uq6-codingqwen38-uq6-thinking
Its model artifacts are the user-requested files from
unsloth/Qwen3.8-27B-GGUF:
Qwen3.8-27B-UD-Q6_K_XL.gguf, 25,299,061,664 bytes;mmproj-BF16.gguf, 931,146,432 bytes;MTP/mtp-Qwen3.8-27B-Q4_0.gguf, 1,369,590,656 bytes.
The requested Hauhau 400K aggregate context did not fit this larger quantization: the MTP context failed a final 1,564 MiB CUDA allocation. A 380K aggregate profile loaded but left only 224 MiB on the most-loaded GPU. The retained maximum operational profile is therefore 348K aggregate, rounded by llama.cpp to 174,080 tokens per slot. It passed two simultaneous public completions with Q8 K/V cache, MTP depth 3, and 1,044 MiB remaining on the most-loaded GPU. The smoke observed 90.7-92.0% draft acceptance. This is a demonstrated operating point, not a long-context saturation proof.
Hauhau retains 400K configured aggregate context, two slots, Q8 K/V cache, and MTP depth 3. FP8 retains 155K max model length, TP2, two sequences, FP8 KV, and three speculative MTP tokens.
Templates and provenance
Three exact templates are archived:
sharp-v22.4.0.jinja, SHA256180e7015759b2b6b57574d6c2ca5c2d19eb2b05a4aaffa80866f71eb1a1fad1a;original-hauhau-fp8.jinja, shared by Hauhau and FP8, SHA256c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041;original-unsloth.jinja, extracted from GGUF metadata, SHA25612827f24b742ea4e80cdc12dbcf9622227056b9f797252a3149263d4f9aaadce.
The Sharp upstream verifier passed base-template, fast-mode, render, minified-roundtrip, tool, and
reasoning-effort checks. extract_gguf_chat_template.py
reproduces the Unsloth extraction when gguf-py is available. Model and template hashes are in
model-artifacts.sha256 and
template-artifacts.sha256; model weights are intentionally not
committed.
Reproduction
Local report server
The report can be served on all host interfaces for access from the LAN or Tailscale network:
./serve_report.sh 8765
The launcher defaults to 0.0.0.0; set REPORT_HOST=localhost when loopback-only access is
desired. On the benchmark host the report is reachable at private workstation service
on the LAN and private workstation service over Tailscale.
Regenerate and verify the canonical portable report with the installed Data Analytics plugin:
node "$DATA_ANALYTICS_PLUGIN_ROOT/skills/build-report/scripts/deliver_portable_artifact.mjs" \
--input artifact.json \
--output report.html
On this benchmark host the enhanced-reader Chromium probe currently remains in its semantic fallback state. The committed report therefore uses the same canonical portable builder with seven embedded light/dark SVG charts, followed by exact structural verification:
node scripts/build_report_static.mjs \
artifact.json report.html \
"$DATA_ANALYTICS_PLUGIN_ROOT/skills/build-report/scripts/build_portable_artifact.mjs"
The generated fallback is self-contained and remains readable without JavaScript. The canonical artifact-to-HTML payload equality check passes; only the enhanced-reader interaction smoke is unavailable on this host.
Recalculate and verify aggregate versus per-response throughput directly from all 1,200 frozen responses:
python3 scripts/audit_throughput_metrics.py --artifact artifact.json
Prepare or reuse the frozen suite, then run an arm:
python3 scripts/qwen38_proxy_accuracy_eval.py run \
--suite data/suite.json \
--model qwen38-uq6-thinking \
--label sharp_unsloth_q6_medium \
--output /tmp/sharp_unsloth_q6_medium.json \
--max-tokens 16000 \
--reasoning-effort medium
Compare two completed arms:
python3 scripts/qwen38_proxy_accuracy_eval.py score \
--left data/original_unsloth_q6_medium.json \
--right data/sharp_unsloth_q6_medium.json \
--output /tmp/unsloth_original_vs_sharp_medium.json
The launcher copies document the exact llama.cpp and vLLM flags used. They retain workstation paths and require the corresponding local builds and model files.
Limits
- This is one deterministic-seed run per arm, not a multi-seed estimate.
- The 100-item sample is useful for paired regression testing but too small to establish the observed one- and two-point differences as general accuracy improvements.
- Aggregate tok/s is total completion output divided by full-suite wall time at concurrency 2. It is server throughput, not the speed of one response.
- Median response tok/s is calculated per response and includes TTFT, batching contention, and generation. It is not a decode-only kernel benchmark.
- Sharp adds prompt instructions, so its prompt-token counts are intentionally higher. Its faster completion in these runs came mostly from shorter outputs, not universally faster token decode.
- The suite scores final answers, not reasoning style, tool use, safety, long-context recall, or vision correctness.