Section 079 · Chapter 10, Anti-Patterns That Create False Confidence
Anti-Patterns: The Aggregate Score Trap
Overall quality can improve while important users, languages, tasks, or risk categories get worse.
RAGaggregate scoreaggregate score trap
What to do
- Do not create so many slices that every result becomes noise.
- Choose the ones that matter, then ensure they have enough sample size or targeted evidence.
- Define slice thresholds before the run.
- Use confidence intervals per slice, risk-weighted reporting, and minimum quality bars for groups where failure has high cost.
Evidence to preserve
- Include protected classes when relevant, regulatory categories, high-value workflows, high-risk actions, and historically weak segments.
- Preserve the inputs, versions, configurations, raw outcomes, and results for RAG, aggregate score, aggregate score trap needed to reproduce work on Anti-Patterns: The Aggregate Score Trap.
- Report results for RAG, aggregate score, aggregate score trap by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Define slice thresholds before the run. Use confidence intervals per slice, risk-weighted reporting, and minimum quality bars for groups where failure has high cost.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Anti-Patterns: The Aggregate Score Trap." Testing AI Knowledge Edition, section 79.
https://jarbon.ai/testing-ai/knowledge/ch079-aggregate-score-trap.html