Section 079 · Chapter 10, Anti-Patterns That Create False Confidence

Anti-Patterns: The Aggregate Score Trap

Overall quality can improve while important users, languages, tasks, or risk categories get worse.

RAGaggregate scoreaggregate score trap

What to do

  1. Do not create so many slices that every result becomes noise.
  2. Choose the ones that matter, then ensure they have enough sample size or targeted evidence.
  3. Define slice thresholds before the run.
  4. Use confidence intervals per slice, risk-weighted reporting, and minimum quality bars for groups where failure has high cost.

Evidence to preserve

  • Include protected classes when relevant, regulatory categories, high-value workflows, high-risk actions, and historically weak segments.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for RAG, aggregate score, aggregate score trap needed to reproduce work on Anti-Patterns: The Aggregate Score Trap.
  • Report results for RAG, aggregate score, aggregate score trap by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Define slice thresholds before the run. Use confidence intervals per slice, risk-weighted reporting, and minimum quality bars for groups where failure has high cost.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Anti-Patterns: The Aggregate Score Trap." Testing AI Knowledge Edition, section 79.

https://jarbon.ai/testing-ai/knowledge/ch079-aggregate-score-trap.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400