Section 026 · Chapter 4, Statistical Tests for AI Quality

Multiple Comparisons and False Discoveries

The more slices, variants, and metrics you inspect, the more likely one lucky result will look real.

multiple comparisonsmultiple comparisons false discoveries

What to do

  1. Run one test and that risk may be acceptable.
  2. Run a hundred tests and the chance that at least one looks significant by luck can become large.
  3. Use holdout sets, preregistered primary metrics, adjusted thresholds, or false-discovery-rate methods when many comparisons are part of the process.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for multiple comparisons, multiple comparisons false discoveries needed to reproduce work on Multiple Comparisons and False Discoveries.
  • Report results for multiple comparisons, multiple comparisons false discoveries by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Separate exploratory analysis, confirmatory analysis, and monitoring. Use holdout sets, preregistered primary metrics, adjusted thresholds, or false-discovery-rate methods when many comparisons are part of the process. When a dashboard contains dozens of segments, report how many comparisons were inspected and which ones were planned before the run.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Multiple Comparisons and False Discoveries." Testing AI Knowledge Edition, section 26.

https://jarbon.ai/testing-ai/knowledge/ch026-multiple-comparisons-false-discoveries.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400