Section 170 · Chapter 20, The Practical Playbook
Measurement Infrastructure Must Know About Variance
If the measurement system ignores its own variance, it will eventually promote lucky noise as product improvement.
What to do
- Use predeclared stopping rules, fresh holdouts, bootstrap intervals, sequential testing discipline, and run logs that preserve every attempt.
- Define runnable checks that exercise variance, sample size, and quality metric.
- Set acceptable outcomes and blocker failures for variance, sample size, and quality metric before running the evaluation.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for variance, sample size, quality metric, RAG needed to reproduce work on Measurement Infrastructure Must Know About Variance.
- Report results for variance, sample size, quality metric, RAG by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
In production work, treat the measurement system as part of the experiment. Model rater variance, sample variance, judge variance, temporal drift, repeated testing, and multiple-comparison effects. Use predeclared stopping rules, fresh holdouts, bootstrap intervals, sequential testing discipline, and run logs that preserve every attempt. A release metric should answer whether the system improved beyond the known noise of both the product and the measuring instrument.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Measurement Infrastructure Must Know About Variance." Testing AI Knowledge Edition, section 170.
https://jarbon.ai/testing-ai/knowledge/ch170-measurement-infrastructure-must-know-variance.html