Section 167 · Chapter 20, The Practical Playbook
Failure Taxonomy for AI Systems
A shared failure language helps teams cluster problems instead of drowning in disconnected bug reports.
RAGconfidence engineerfailure taxonomy
What to do
- Start with factual failures: wrong facts, invented facts, stale facts, missing required facts, or unsupported claims.
- Define runnable checks that exercise RAG, confidence engineer, and failure taxonomy.
- Set acceptable outcomes and blocker failures for RAG, confidence engineer, and failure taxonomy before running the evaluation.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for RAG, confidence engineer, failure taxonomy needed to reproduce work on Failure Taxonomy for AI Systems.
- Report results for RAG, confidence engineer, failure taxonomy by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Failure taxonomy should connect to severity, affected slices, root-cause hypotheses, owners, regression cases, and incident metrics.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Failure Taxonomy for AI Systems." Testing AI Knowledge Edition, section 167.
https://jarbon.ai/testing-ai/knowledge/ch167-failure-taxonomy.html