Back to the experimentSupporting notebook

The Chat Template That Cut Qwen3.8 Completion Tokens by 16–21%

Source: qwen38/templates/2026-08-27-medium-xhigh-template-ab/BLOG.md · revision 6550ead3945b

Qwen Sharp v22.4.0 made every tested Qwen3.8 deployment more concise, but the quality result depended on the weights. On two Q6 GGUF models it preserved or improved combined accuracy; on FP8 it traded two correct answers for shorter output.

Executive summary

Why test a chat template at all?

A chat template is often treated like formatting glue: it turns message objects into the token sequence a model reads. In a thinking model, however, that sequence also carries instructions about reasoning effort, answer structure, tools, and when to stop. A template can therefore change how much reasoning the model emits, how quickly a request finishes, and whether the final answer remains parseable.

We wanted to know whether Qwen Sharp Chat Templates v22.4.0 improved a real two-GPU Qwen3.8 deployment—not only whether it rendered correctly. The decision had three parts:

  1. Did Sharp change final-answer correctness?
  2. Did it reduce completion length and end-to-end time?
  3. Did those effects hold across different Qwen3.8 weights and runtimes?

The tested deployments were HauhauCS Aggressive Q6_K_P GGUF, official Unsloth UD-Q6_K_XL GGUF, and OrcaRouter uncensored FP8. The two GGUF models ran through llama.cpp FastMTP; FP8 ran through vLLM MTP. Every deployment used two RTX 3090 GPUs, speculative depth 3, and maximum concurrency 2.

The experiment

The frozen suite contained 100 questions: 40 MMLU-Pro, 20 ARC-Challenge, 20 HellaSwag, and 20 GSM8K. Each model ran four arms:

That produced 12 complete runs and 1,200 scored responses. Every arm allowed up to 16,000 completion tokens. Multiple-choice answers had to end with FINAL: <LETTER> and GSM8K answers with FINAL: <NUMBER>. Unparsed or length-capped responses remained incorrect; they were not manually rescued.

The public thinking routes used the same effective sampling parameters in every arm: temperature 1.0, top-p 0.95, top-k 20, presence penalty 1.5, and min-p 0.0. The test runner used two worker threads and fixed source rows and seeds.

The key formulas were:

completion-token reduction = (stock completion tokens - Sharp completion tokens) / stock completion tokens

net total-token reduction = (stock total tokens - Sharp total tokens) / stock total tokens

aggregate tok/s = all completion tokens / complete 100-item run wall seconds

response tok/s = one response's completion tokens / that response's end-to-end wall seconds

Aggregate tok/s measures server capacity with two concurrent requests. For user-visible speed, we report the median of the 100 individual response rates separately. Neither metric is a pure warm decode-only measurement because response wall time includes TTFT and batching effects.

Sharp consistently reduced output length

Across medium and xhigh together, the result was consistent: Sharp produced fewer completion tokens on every model.

Model Stock accuracy Sharp accuracy Stock completion tokens Sharp completion tokens Completion reduction Net total-token reduction Wall-time reduction
Hauhau Q6 179/200 180/200 295,356 234,586 20.6% 9.0% 24.7%
Unsloth Q6 178/200 180/200 245,289 196,041 20.1% 6.5% 19.1%
FP8 178/200 176/200 249,214 209,229 16.0% 3.2% 11.9%

Figure 1 in the interactive report plots the six matched completion-token reductions. The source is the completion_tokens field in the 12 frozen run JSON files. Alt text: Sharp reduced completion tokens in all six matched tests, with reductions ranging from 9.8% to 28.6%.

The combined figures hide an important difference between medium and xhigh:

Model Effort Sharp accuracy delta Completion tokens saved Completion reduction Wall-time reduction
Hauhau Q6 medium -1 25,173 28.2% 23.9%
Hauhau Q6 xhigh +2 35,597 17.3% 25.0%
Unsloth Q6 medium 0 11,327 14.8% 11.5%
Unsloth Q6 xhigh +2 37,921 22.5% 22.0%
FP8 medium 0 23,724 28.6% 21.9%
FP8 xhigh -2 16,261 9.8% 8.2%

On Hauhau and Unsloth xhigh, Sharp was both shorter and two answers better. On FP8 xhigh, it was shorter but two answers worse. Medium FP8 is the clearest efficiency result: Sharp matched stock at 88/100 while cutting completion output by 28.6% and wall time by 21.9%.

Shorter completion output is not the same as fewer total tokens

Sharp added 15,500 prompt tokens across each 100-item run—about 155 extra prompt tokens per question. That overhead matters.

For example, Unsloth medium dropped from 76,516 to 65,189 completion tokens, a 14.8% reduction. But prompt tokens rose from 18,784 to 34,284. Total prompt-plus-completion usage therefore increased from 95,300 to 99,473, or 4.4%, in that particular arm.

Xhigh Unsloth went the other direction: its completion reduction was large enough to overcome the prompt overhead, reducing total tokens by 11.9%. Across both efforts, Unsloth still achieved a net 6.5% total-token reduction.

This distinction is why the report exposes prompt tokens, completion tokens, total tokens, and tokens per response separately. “The model answered more concisely” is directly observed. “The deployment used 20% fewer billable tokens” would be inaccurate unless only completion tokens were billed.

What response length looked like

The average completion length makes the effect tangible:

Model Stock medium Sharp medium Stock xhigh Sharp xhigh
Hauhau Q6 892 tokens 640 tokens 2,061 tokens 1,706 tokens
Unsloth Q6 765 tokens 652 tokens 1,688 tokens 1,309 tokens
FP8 830 tokens 593 tokens 1,662 tokens 1,500 tokens

Xhigh did not automatically buy better answers. Unsloth Sharp scored 90/100 at both efforts, but xhigh emitted roughly twice as many completion tokens per response and took 2.36 times as long. For that model, medium is the obvious operating point.

Hauhau Sharp was different: xhigh improved from 89/100 to the experiment-high 91/100, at the cost of moving from 640 to 1,706 tokens per answer. That profile makes sense only when the extra observed correctness matters more than latency and output length.

FP8 was still much faster per response

FP8’s median individual response rate ranged from 52.6 to 68.3 tok/s. Hauhau and Unsloth Q6 ranged from 30.2 to 37.2 tok/s. These are the numbers to compare with the speed visible to a user waiting for one response.

FP8 also remained the aggregate throughput leader. With concurrency 2 it ranged from 97.3 to 134.7 tok/s, while Q6 ranged from 60.7 to 74.1 tok/s. At medium effort, FP8 stock reached 68.3 median response tok/s and 134.7 aggregate tok/s; Sharp reached 61.7 and 123.1 respectively. Sharp nevertheless completed the suite sooner because it generated fewer tokens: 481.4 versus 616.2 seconds.

The interactive report now uses separate charts for aggregate throughput and median per-response rate. Alt text: FP8 has substantially higher individual and aggregate token rates than either Q6 deployment, while Sharp generally shortens total run time by reducing the amount of output.

This is a useful operational lesson: a lower tok/s value does not necessarily mean the user waits longer. If a template elicits a shorter answer, the request can finish sooner even when the backend processes fewer completion tokens per second.

Why Sharp may be shortening responses

The measured mechanism is straightforward: Sharp adds more prompt tokens and the model emits fewer completion tokens. The exact causal mechanism was not isolated.

A plausible explanation is that Sharp’s explicit reasoning-effort and answer-shaping instructions provide a stronger stopping and response contract. That may reduce meandering or repeated reasoning. The current benchmark did not classify reasoning styles or compare semantic redundancy, so this remains a hypothesis rather than a validated explanation.

The FP8 result is also a reminder that template behavior depends on the weights and runtime. A template that improved the two Q6 variants did not preserve xhigh accuracy on the FP8 model. Compatibility at render time is not enough; each model needs its own outcome test.

What we recommend

For a balanced Q6 default, use Unsloth UD-Q6_K_XL with Sharp at medium effort. It scored 90/100, averaged 652 completion tokens per response, delivered 34.8 median response tok/s and 71.3 aggregate tok/s, and finished in 913.8 seconds. Xhigh produced no accuracy gain for this model.

For the highest observed score, use Hauhau Q6 with Sharp at xhigh. It reached 91/100, but its 1,706-token average response and 2,546.9-second suite time make it a quality-first route rather than the everyday default.

For the deployed everyday route, use FP8 with Sharp at medium effort. It matched stock medium at 88/100 while using 28.6% fewer completion tokens and finishing 21.9% sooner. Stock xhigh remains the highest-scoring FP8 arm at 90/100, but that two-answer difference was not statistically significant and xhigh more than tripled the Sharp-medium suite time.

What this does not prove

Each arm was run once with one deterministic seed over a 100-item suite. The paired accuracy differences were small: every exact McNemar p-value was at least 0.50, and every paired bootstrap interval included zero. The experiment therefore supports regression and deployment choices, not a universal ranking.

The suite evaluates final answers, not long-context recall, tool use, vision, safety, writing quality, or conversational preference. End-to-end wall time may include cold model loading or first-shape JIT. Aggregate and response tok/s are not warm decode-only measurements. The next controlled experiment should repeat the suite across multiple seeds and add a warm, fixed-output-length decode benchmark to isolate raw backend speed from response-length behavior.

Reproduce or audit the result

The complete raw outputs, frozen suite, exact templates, model hashes, paired score reports, and runner are stored alongside this post. A representative run is:

python3 scripts/qwen38_proxy_accuracy_eval.py run \
  --suite data/suite.json \
  --model qwen38-uq6-thinking \
  --label sharp_unsloth_q6_medium \
  --output /tmp/sharp_unsloth_q6_medium.json \
  --max-tokens 16000 \
  --reasoning-effort medium

Use README.md for the full methodology, artifact.json for the chart-ready reviewed snapshot, and report.html for the interactive comparison.

Evidence appendix

Glossary: “stock” means the model’s original chat template; “Sharp” means Qwen Sharp v22.4.0; “medium” and “xhigh” are explicit reasoning-effort values; “MTP” is multi-token prediction used for speculative decoding; “aggregate tok/s” sums both concurrent streams; “median response tok/s” is the median of the 100 per-response end-to-end completion rates.