Section 045 · Chapter 6, Building Evals That Matter
Benchmarking: Quality Is Relative Now
When absolute truth is hard to measure, relative quality can still tell you whether you are competitive, broken, unusual, or missing something obvious.
What to do
- Do not pretend the comparison is perfectly controlled.
- Report ranks, percentiles, confidence intervals when available, and qualitative gaps.
- Separate "competitors do this" from "users need this" and "regulators require this." Benchmarking is evidence, not authority.
Evidence to preserve
- Report ranks, percentiles, confidence intervals when available, and qualitative gaps.
- Preserve the inputs, versions, configurations, raw outcomes, and results for benchmark, benchmarking quality relative needed to reproduce work on Benchmarking: Quality Is Relative Now.
- Report results for benchmark, benchmarking quality relative by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
The deeper move is to make benchmark design explicit: define the peer set, task set, sampling frame, measurement method, normalizations, and known unfairness. Competitors may differ in traffic, geography, device mix, business model, legal obligations, data access, and product goals. Do not pretend the comparison is perfectly controlled.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Benchmarking: Quality Is Relative Now." Testing AI Knowledge Edition, section 45.
https://jarbon.ai/testing-ai/knowledge/ch045-benchmarking-quality-relative.html