Section 011 · Chapter 2, From Tests to Release Evidence
Rare Failure Hunting
Average quality can look excellent while rare catastrophic failures still make the system unsafe.
rare failureRAGrare failure hunting
What to do
- Track the worst observed output, the number of critical failures, the percentage below a threshold, safety failure rate, privacy failure rate, and policy violation rate.
- Define runnable checks that exercise rare failure, RAG, and rare failure hunting.
- Set acceptable outcomes and blocker failures for rare failure, RAG, and rare failure hunting before running the evaluation.
Evidence to preserve
- Track the worst observed output, the number of critical failures, the percentage below a threshold, safety failure rate, privacy failure rate, and policy violation rate.
- Include examples of the worst outputs in the report so decision-makers can see the risk directly.
- Preserve the inputs, versions, configurations, raw outcomes, and results for rare failure, RAG, rare failure hunting needed to reproduce work on Rare Failure Hunting.
- Report results for rare failure, RAG, rare failure hunting by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Expert rare-failure work treats zero observed failures carefully. If you test 100 cases and see zero failures, you have evidence, not proof. The upper bound on the plausible failure rate may still be too high for safety-critical behavior.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Rare Failure Hunting." Testing AI Knowledge Edition, section 11.
https://jarbon.ai/testing-ai/knowledge/ch011-rare-failure-hunting.html