Back to the experimentSupporting notebook

Chart map

Source: qwen38/templates/2026-08-27-medium-xhigh-template-ab/CHART_MAP.md · revision 6550ead3945b

Report section Question Chart Fields Supported takeaway Palette
Accuracy Which run answered the most items correctly? Comparison bar run, correct Best observed score was Hauhau Sharp xhigh at 91/100 Single blue root; direct run labels
Latency Which run completed fastest? Comparison bar run, wall_seconds FP8 completed fastest; Sharp reduced wall time in all matched arms Single blue root; lower is better in subtitle
Response length How long were the answers? Comparison bar run, avg_completion_tokens Sharp shortened average completion length in all six matched comparisons Single gold root; exact tooltip totals
Aggregate throughput What two-request server throughput was measured? Comparison bar run, aggregate_tps FP8 produced roughly 1.5–1.9x the Q6 aggregate throughput Single blue root; concurrency and metric boundary in subtitle
Per-response rate What completion rate did one response typically experience? Comparison bar run, median_response_tps FP8 medians were 52.6–68.3 tok/s versus 30.2–37.2 tok/s for Q6 Single olive root; end-to-end response definition in subtitle
Token savings How much completion output did Sharp remove? Comparison bar comparison, token_savings_pct Sharp reduced completion output by 9.8–28.6% in every matched arm Single olive root; exact tokens saved in tooltip
Trade-off How do accuracy and end-to-end time relate? Scatter wall_seconds, correct, model Higher effort often costs much more time for small accuracy movement Model color plus direct run tooltip

All chart datasets are bounded snapshots in artifact.json and trace to the 12 raw run JSON files under data/. The static semantic fallback exposes the same reviewed rows when the enhanced reader is unavailable.