Section 029 · Chapter 5, Judges, Humans, and Disagreement
Human Calibration of LLM Judges
Before relying on an LLM judge at scale, builders need to know whether it scores like a trusted human reviewer.
LLM judgehuman calibrationhuman reviewhuman calibration llm judges
What to do
- Compare it against the old judge.
- Check for overfitting to the examples.
- Watch slices where the fine-tuned judge becomes more confident but less correct.
- Treat judge quality as something you test, not something you assume.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for LLM judge, human calibration, human review, human calibration llm judges needed to reproduce work on Human Calibration of LLM Judges.
- Report results for LLM judge, human calibration, human review, human calibration llm judges by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
In production work, track judge-human agreement over time, by category, and by severity. A judge can be acceptable for low-risk style checks and unacceptable for regulated policy decisions. Calibration should produce routing rules, not just a single accuracy number.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Human Calibration of LLM Judges." Testing AI Knowledge Edition, section 29.
https://jarbon.ai/testing-ai/knowledge/ch029-human-calibration-llm-judges.html