Section 033 · Chapter 5, Judges, Humans, and Disagreement
Using Raters Well
Human raters are not a checkbox. They are an evaluation instrument that needs selection, calibration, workflow design, and quality control.
exact assertionsLLM judgehuman raterraters
What to do
- Use overlap deliberately.
- Keep raters blind when possible.
- Track disagreement as data.
- Track inter-rater agreement, rater-specific bias, calibration drift, fatigue effects, adjudication outcomes, and whether the rater population matches the user population.
Evidence to preserve
- Track disagreement as data.
- Track inter-rater agreement, rater-specific bias, calibration drift, fatigue effects, adjudication outcomes, and whether the rater population matches the user population.
- Preserve the inputs, versions, configurations, raw outcomes, and results for exact assertions, LLM judge, human rater, raters needed to reproduce work on Using Raters Well.
- Report results for exact assertions, LLM judge, human rater, raters by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
At scale, treat raters as measurement instruments. Track inter-rater agreement, rater-specific bias, calibration drift, fatigue effects, adjudication outcomes, and whether the rater population matches the user population. If raters and users disagree systematically, the eval is measuring the wrong audience.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Using Raters Well." Testing AI Knowledge Edition, section 33.
https://jarbon.ai/testing-ai/knowledge/ch033-raters.html