Section 021 · Chapter 4, Statistical Tests for AI Quality
Null Hypothesis: What Are We Actually Testing?
The null hypothesis is the boring default: assume the new AI system is not really better until the evidence is strong enough to challenge that assumption.
What to do
- Define runnable checks that exercise null hypothesis and null hypothesis we actually.
- Set acceptable outcomes and blocker failures for null hypothesis and null hypothesis we actually before running the evaluation.
- Run representative cases for null hypothesis and null hypothesis we actually and preserve the failures that would change the decision.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for null hypothesis, null hypothesis we actually needed to reproduce work on Null Hypothesis: What Are We Actually Testing?.
- Report results for null hypothesis, null hypothesis we actually by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
The deeper move is to define the null hypothesis before running the evaluation. The null hypothesis is the boring default claim: nothing meaningful changed, the new model is not better, the new policy did not reduce harm, or the new agent did not improve task completion beyond ordinary noise. An evaluation is the planned measurement used to challenge that claim.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Null Hypothesis: What Are We Actually Testing?." Testing AI Knowledge Edition, section 21.
https://jarbon.ai/testing-ai/knowledge/ch021-null-hypothesis-we-actually.html