Back to the experimentSupporting notebook

Qwen3.8 template and reasoning-effort A/B on the 12434 proxy

Source: qwen38/templates/2026-08-27-medium-xhigh-template-ab/README.md · revision 6550ead3945b

This experiment compares the stock Qwen3.8 chat templates with Qwen Sharp Chat Templates v22.4.0 at medium and xhigh reasoning effort. It covers three live two-GPU deployments:

The run matrix contains 12 complete 100-item runs. Every raw response, reasoning trace, token count, finish reason, paired comparison, suite item, and scoring script is included under data/ and scripts/.

Read the publication-ready blog post or open the self-contained interactive chart report. The updated report includes accuracy, wall time, completion tokens, prompt-plus-completion totals, average response length, median per-response tok/s, aggregate tok/s at concurrency 2, and matched Sharp-versus-stock savings.

Deployment outcome

Correctness was the first selection criterion and observed completion time was the tiebreaker. The retained proxy configuration is:

Deployment Retained template Recommended effort Why
Hauhau Q6 Sharp v22.4.0 xhigh when maximum benchmark correctness matters; medium for latency Sharp scored 180/200 across both efforts versus 179/200 stock and used 24.7% less combined wall time. Its best single run was 91/100 at xhigh.
Unsloth UD-Q6_K_XL Sharp v22.4.0 medium Both Sharp efforts scored 90/100, but medium was 2.36x faster and used about half the completion tokens. Sharp scored 180/200 versus 178/200 stock.
OrcaRouter FP8 Sharp v22.4.0 operational default medium Sharp medium is the deployed speed/token-efficiency choice. Stock xhigh reached 90/100 versus 88/100 for Sharp, but the two-answer difference was not statistically significant.

These one- and two-answer differences are operational choices, not proof of a general quality winner: no paired accuracy contrast was statistically significant in this 100-item sample.

Complete results

All runs used concurrency 2 and a 16,000-token maximum completion allowance. Wall time is end-to-end for the complete 100-item run and can include a cold model load or first-shape JIT.

Model Template Effort Correct Wall seconds Completion tokens Median response tok/s Aggregate tok/s (c=2) Unparsed Length stops
Hauhau Q6 Sharp medium 89/100 918.211 64,036 34.21 69.74 0 0
Hauhau Q6 Sharp xhigh 91/100 2,546.913 170,550 33.77 66.96 2 2
Hauhau Q6 Stock medium 90/100 1,206.307 89,209 36.30 73.95 0 0
Hauhau Q6 Stock xhigh 89/100 3,395.125 206,147 30.79 60.72 4 5
Unsloth Q6 Sharp medium 90/100 913.819 65,189 34.79 71.34 0 0
Unsloth Q6 Sharp xhigh 90/100 2,155.354 130,852 30.22 60.71 1 1
Unsloth Q6 Stock medium 90/100 1,032.878 76,516 37.17 74.08 0 0
Unsloth Q6 Stock xhigh 88/100 2,763.000 168,773 30.57 61.08 3 4
FP8 Sharp medium 88/100 481.396 59,255 61.66 123.09 0 0
FP8 Sharp xhigh 88/100 1,541.898 149,974 52.56 97.27 4 4
FP8 Stock medium 88/100 616.227 82,979 68.26 134.66 0 0
FP8 Stock xhigh 90/100 1,679.790 166,235 54.59 98.96 4 4

All 12 runs had zero transport/backend errors. A response that reached 16K without a parseable FINAL: answer remained wrong; no manual credit was added.

What xhigh changed

Using the retained template for each model:

Model/template Accuracy change, medium to xhigh Wall-time multiplier Completion-token multiplier
Hauhau / Sharp +2 answers 2.77x 2.66x
Unsloth / Sharp 0 answers 2.36x 2.01x
FP8 / Stock +2 answers 2.73x 2.00x

The observed +2 changes for Hauhau and FP8 were not significant: paired bootstrap intervals included zero and exact McNemar p-values were 0.727 and 0.688, respectively. Unsloth Sharp was an accuracy tie with McNemar p=1.0.

Template A/B paired correctness

The table reports Sharp minus stock on the same 100 item IDs.

Model Effort Stock Sharp Stock-only correct Sharp-only correct Difference Exact McNemar p
Hauhau Q6 medium 90 89 1 0 -1 point 1.000
Hauhau Q6 xhigh 89 91 1 3 +2 points 0.625
Unsloth Q6 medium 90 90 2 2 0 points 1.000
Unsloth Q6 xhigh 88 90 1 3 +2 points 0.625
FP8 medium 88 88 2 2 0 points 1.000
FP8 xhigh 90 88 2 0 -2 points 0.500

Correctness protocol

The frozen suite has 100 examples:

Prompts, source row IDs, gold answers, seeds, and source provenance are stored in data/suite.json. Multiple-choice responses had to end with FINAL: <LETTER> and GSM8K responses with FINAL: <NUMBER>. The scorer normalizes basic numeric formatting and uses the final matching answer. It does not judge free-form reasoning or rescue a truncated answer manually.

Both template families explicitly support reasoning_effort=medium|xhigh; the effort value was passed through chat_template_kwargs with thinking enabled and preserved. The public thinking routes forced the same effective sampling preset in every arm: temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5, and min-p 0.0. The runner used two worker threads and fixed seeds.

Runtime profiles and the new Unsloth deployment

The exact retained family definitions are in configs/final-model-profiles.json.

The newly added public route family is:

Its model artifacts are the user-requested files from unsloth/Qwen3.8-27B-GGUF:

The requested Hauhau 400K aggregate context did not fit this larger quantization: the MTP context failed a final 1,564 MiB CUDA allocation. A 380K aggregate profile loaded but left only 224 MiB on the most-loaded GPU. The retained maximum operational profile is therefore 348K aggregate, rounded by llama.cpp to 174,080 tokens per slot. It passed two simultaneous public completions with Q8 K/V cache, MTP depth 3, and 1,044 MiB remaining on the most-loaded GPU. The smoke observed 90.7-92.0% draft acceptance. This is a demonstrated operating point, not a long-context saturation proof.

Hauhau retains 400K configured aggregate context, two slots, Q8 K/V cache, and MTP depth 3. FP8 retains 155K max model length, TP2, two sequences, FP8 KV, and three speculative MTP tokens.

Templates and provenance

Three exact templates are archived:

The Sharp upstream verifier passed base-template, fast-mode, render, minified-roundtrip, tool, and reasoning-effort checks. extract_gguf_chat_template.py reproduces the Unsloth extraction when gguf-py is available. Model and template hashes are in model-artifacts.sha256 and template-artifacts.sha256; model weights are intentionally not committed.

Reproduction

Local report server

The report can be served on all host interfaces for access from the LAN or Tailscale network:

./serve_report.sh 8765

The launcher defaults to 0.0.0.0; set REPORT_HOST=localhost when loopback-only access is desired. On the benchmark host the report is reachable at private workstation service on the LAN and private workstation service over Tailscale.

Regenerate and verify the canonical portable report with the installed Data Analytics plugin:

node "$DATA_ANALYTICS_PLUGIN_ROOT/skills/build-report/scripts/deliver_portable_artifact.mjs" \
  --input artifact.json \
  --output report.html

On this benchmark host the enhanced-reader Chromium probe currently remains in its semantic fallback state. The committed report therefore uses the same canonical portable builder with seven embedded light/dark SVG charts, followed by exact structural verification:

node scripts/build_report_static.mjs \
  artifact.json report.html \
  "$DATA_ANALYTICS_PLUGIN_ROOT/skills/build-report/scripts/build_portable_artifact.mjs"

The generated fallback is self-contained and remains readable without JavaScript. The canonical artifact-to-HTML payload equality check passes; only the enhanced-reader interaction smoke is unavailable on this host.

Recalculate and verify aggregate versus per-response throughput directly from all 1,200 frozen responses:

python3 scripts/audit_throughput_metrics.py --artifact artifact.json

Prepare or reuse the frozen suite, then run an arm:

python3 scripts/qwen38_proxy_accuracy_eval.py run \
  --suite data/suite.json \
  --model qwen38-uq6-thinking \
  --label sharp_unsloth_q6_medium \
  --output /tmp/sharp_unsloth_q6_medium.json \
  --max-tokens 16000 \
  --reasoning-effort medium

Compare two completed arms:

python3 scripts/qwen38_proxy_accuracy_eval.py score \
  --left data/original_unsloth_q6_medium.json \
  --right data/sharp_unsloth_q6_medium.json \
  --output /tmp/unsloth_original_vs_sharp_medium.json

The launcher copies document the exact llama.cpp and vLLM flags used. They retain workstation paths and require the corresponding local builds and model files.

Limits