Section 026 · Chapter 4, Statistical Tests for AI Quality
Multiple Comparisons and False Discoveries
The more slices, variants, and metrics you inspect, the more likely one lucky result will look real.
What to do
- Run one test and that risk may be acceptable.
- Run a hundred tests and the chance that at least one looks significant by luck can become large.
- Use holdout sets, preregistered primary metrics, adjusted thresholds, or false-discovery-rate methods when many comparisons are part of the process.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for multiple comparisons, multiple comparisons false discoveries needed to reproduce work on Multiple Comparisons and False Discoveries.
- Report results for multiple comparisons, multiple comparisons false discoveries by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Separate exploratory analysis, confirmatory analysis, and monitoring. Use holdout sets, preregistered primary metrics, adjusted thresholds, or false-discovery-rate methods when many comparisons are part of the process. When a dashboard contains dozens of segments, report how many comparisons were inspected and which ones were planned before the run.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Multiple Comparisons and False Discoveries." Testing AI Knowledge Edition, section 26.
https://jarbon.ai/testing-ai/knowledge/ch026-multiple-comparisons-false-discoveries.html