Section 168 · Chapter 20, The Practical Playbook
AI Always Fails
The useful question is not whether AI will fail. It is where, how often, how badly, and whether you already know which inputs are likely to break it.
What to do
- Build a useful failure map instead of promising a perfect AI system.
- Start with domain experts.
- Ask what users misunderstand, what policies are subtle, which cases are rare but severe, and which inputs even humans find difficult.
- Treat failure discovery as a continuous measurement problem.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for generated code, personalization, always fails needed to reproduce work on AI Always Fails.
- Report results for generated code, personalization, always fails by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Treat failure discovery as a continuous measurement problem. Combine production trace mining, synthetic edge-case generation, adversarial testing, human review, clustering, severity scoring, and slice-level confidence intervals. The output should be a failure taxonomy with owners, detection signals, regression cases, escalation rules, and release thresholds.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "AI Always Fails." Testing AI Knowledge Edition, section 168.
https://jarbon.ai/testing-ai/knowledge/ch168-always-fails.html