Section 167 · Chapter 20, The Practical Playbook

Failure Taxonomy for AI Systems

A shared failure language helps teams cluster problems instead of drowning in disconnected bug reports.

RAGconfidence engineerfailure taxonomy

What to do

  1. Start with factual failures: wrong facts, invented facts, stale facts, missing required facts, or unsupported claims.
  2. Define runnable checks that exercise RAG, confidence engineer, and failure taxonomy.
  3. Set acceptable outcomes and blocker failures for RAG, confidence engineer, and failure taxonomy before running the evaluation.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for RAG, confidence engineer, failure taxonomy needed to reproduce work on Failure Taxonomy for AI Systems.
  • Report results for RAG, confidence engineer, failure taxonomy by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Failure taxonomy should connect to severity, affected slices, root-cause hypotheses, owners, regression cases, and incident metrics.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Failure Taxonomy for AI Systems." Testing AI Knowledge Edition, section 167.

https://jarbon.ai/testing-ai/knowledge/ch167-failure-taxonomy.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400