Section 029 · Chapter 5, Judges, Humans, and Disagreement

Human Calibration of LLM Judges

Before relying on an LLM judge at scale, builders need to know whether it scores like a trusted human reviewer.

LLM judgehuman calibrationhuman reviewhuman calibration llm judges

What to do

  1. Compare it against the old judge.
  2. Check for overfitting to the examples.
  3. Watch slices where the fine-tuned judge becomes more confident but less correct.
  4. Treat judge quality as something you test, not something you assume.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for LLM judge, human calibration, human review, human calibration llm judges needed to reproduce work on Human Calibration of LLM Judges.
  • Report results for LLM judge, human calibration, human review, human calibration llm judges by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

In production work, track judge-human agreement over time, by category, and by severity. A judge can be acceptable for low-risk style checks and unacceptable for regulated policy decisions. Calibration should produce routing rules, not just a single accuracy number.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Human Calibration of LLM Judges." Testing AI Knowledge Edition, section 29.

https://jarbon.ai/testing-ai/knowledge/ch029-human-calibration-llm-judges.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400