Section 081 · Chapter 10, Anti-Patterns That Create False Confidence

Anti-Patterns: Treating the Judge as Truth

LLM judges are useful evaluators, not objective measurement devices handed down from the sky.

varianceLLM judgerubrictreating judge truth

What to do

  1. Compare judge scores to human raters on a representative sample.
  2. Track judge drift when the model changes.
  3. Compare LLM judges against humans.
  4. Compare humans against each other.
  5. Review severe failures manually.

Evidence to preserve

  • Track judge drift when the model changes.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for variance, LLM judge, rubric, treating judge truth needed to reproduce work on Anti-Patterns: Treating the Judge as Truth.
  • Report results for variance, LLM judge, rubric, treating judge truth by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

In production work, track judge-human agreement, judge variance, position bias, rubric sensitivity, model-version drift, and category-specific reliability. Treat judge output as evidence with uncertainty.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Anti-Patterns: Treating the Judge as Truth." Testing AI Knowledge Edition, section 81.

https://jarbon.ai/testing-ai/knowledge/ch081-treating-judge-truth.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400