| Report section | Question | Chart | Fields | Supported takeaway | Palette |
|---|---|---|---|---|---|
| Accuracy | Which run answered the most items correctly? | Comparison bar | run, correct |
Best observed score was Hauhau Sharp xhigh at 91/100 | Single blue root; direct run labels |
| Latency | Which run completed fastest? | Comparison bar | run, wall_seconds |
FP8 completed fastest; Sharp reduced wall time in all matched arms | Single blue root; lower is better in subtitle |
| Response length | How long were the answers? | Comparison bar | run, avg_completion_tokens |
Sharp shortened average completion length in all six matched comparisons | Single gold root; exact tooltip totals |
| Aggregate throughput | What two-request server throughput was measured? | Comparison bar | run, aggregate_tps |
FP8 produced roughly 1.5–1.9x the Q6 aggregate throughput | Single blue root; concurrency and metric boundary in subtitle |
| Per-response rate | What completion rate did one response typically experience? | Comparison bar | run, median_response_tps |
FP8 medians were 52.6–68.3 tok/s versus 30.2–37.2 tok/s for Q6 | Single olive root; end-to-end response definition in subtitle |
| Token savings | How much completion output did Sharp remove? | Comparison bar | comparison, token_savings_pct |
Sharp reduced completion output by 9.8–28.6% in every matched arm | Single olive root; exact tokens saved in tooltip |
| Trade-off | How do accuracy and end-to-end time relate? | Scatter | wall_seconds, correct, model |
Higher effort often costs much more time for small accuracy movement | Model color plus direct run tooltip |
All chart datasets are bounded snapshots in artifact.json and trace to the 12 raw run JSON files
under data/. The static semantic fallback exposes the same reviewed rows when the enhanced reader
is unavailable.