Section 081 · Chapter 10, Anti-Patterns That Create False Confidence
Anti-Patterns: Treating the Judge as Truth
LLM judges are useful evaluators, not objective measurement devices handed down from the sky.
varianceLLM judgerubrictreating judge truth
What to do
- Compare judge scores to human raters on a representative sample.
- Track judge drift when the model changes.
- Compare LLM judges against humans.
- Compare humans against each other.
- Review severe failures manually.
Evidence to preserve
- Track judge drift when the model changes.
- Preserve the inputs, versions, configurations, raw outcomes, and results for variance, LLM judge, rubric, treating judge truth needed to reproduce work on Anti-Patterns: Treating the Judge as Truth.
- Report results for variance, LLM judge, rubric, treating judge truth by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
In production work, track judge-human agreement, judge variance, position bias, rubric sensitivity, model-version drift, and category-specific reliability. Treat judge output as evidence with uncertainty.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Anti-Patterns: Treating the Judge as Truth." Testing AI Knowledge Edition, section 81.
https://jarbon.ai/testing-ai/knowledge/ch081-treating-judge-truth.html