Section 025 · Chapter 4, Statistical Tests for AI Quality
Power Analysis and Minimum Detectable Effect
Before asking whether a change won, builders should decide what size of win would actually matter.
sample sizepower analysispower analysis minimum detectable effect
What to do
- Define runnable checks that exercise sample size, power analysis, and power analysis minimum detectable effect.
- Set acceptable outcomes and blocker failures for sample size, power analysis, and power analysis minimum detectable effect before running the evaluation.
- Run representative cases for sample size, power analysis, and power analysis minimum detectable effect and preserve the failures that would change the decision.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for sample size, power analysis, power analysis minimum detectable effect needed to reproduce work on Power Analysis and Minimum Detectable Effect.
- Report results for sample size, power analysis, power analysis minimum detectable effect by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Expert teams distinguish statistical power from business value. High power helps detect a chosen effect, but the minimum meaningful effect should come from product risk, user impact, cost, and operational tradeoffs.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Power Analysis and Minimum Detectable Effect." Testing AI Knowledge Edition, section 25.
https://jarbon.ai/testing-ai/knowledge/ch025-power-analysis-minimum-detectable-effect.html