Section 030 · Chapter 5, Judges, Humans, and Disagreement

Inter-Rater Agreement

When reviewers disagree often, the evaluation system may need as much attention as the product being evaluated.

evaluation criteriaLLM judgeinter-rater agreementrubricattentioninter rater agreement

What to do

  1. Define runnable checks that exercise evaluation criteria, LLM judge, and inter-rater agreement.
  2. Set acceptable outcomes and blocker failures for evaluation criteria, LLM judge, and inter-rater agreement before running the evaluation.
  3. Run representative cases for evaluation criteria, LLM judge, and inter-rater agreement and preserve the failures that would change the decision.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for evaluation criteria, LLM judge, inter-rater agreement, rubric needed to reproduce work on Inter-Rater Agreement.
  • Report results for evaluation criteria, LLM judge, inter-rater agreement, rubric by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Expert teams use agreement statistics carefully. Cohen's kappa, Fleiss' kappa, and Krippendorff's alpha adjust for chance agreement, but they still depend on label design, prevalence, reviewer training, and whether the task is ordinal or categorical.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Inter-Rater Agreement." Testing AI Knowledge Edition, section 30.

https://jarbon.ai/testing-ai/knowledge/ch030-inter-rater-agreement.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400