Section 030 · Chapter 5, Judges, Humans, and Disagreement
Inter-Rater Agreement
When reviewers disagree often, the evaluation system may need as much attention as the product being evaluated.
evaluation criteriaLLM judgeinter-rater agreementrubricattentioninter rater agreement
What to do
- Define runnable checks that exercise evaluation criteria, LLM judge, and inter-rater agreement.
- Set acceptable outcomes and blocker failures for evaluation criteria, LLM judge, and inter-rater agreement before running the evaluation.
- Run representative cases for evaluation criteria, LLM judge, and inter-rater agreement and preserve the failures that would change the decision.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for evaluation criteria, LLM judge, inter-rater agreement, rubric needed to reproduce work on Inter-Rater Agreement.
- Report results for evaluation criteria, LLM judge, inter-rater agreement, rubric by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Expert teams use agreement statistics carefully. Cohen's kappa, Fleiss' kappa, and Krippendorff's alpha adjust for chance agreement, but they still depend on label design, prevalence, reviewer training, and whether the task is ordinal or categorical.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Inter-Rater Agreement." Testing AI Knowledge Edition, section 30.
https://jarbon.ai/testing-ai/knowledge/ch030-inter-rater-agreement.html