Section 031 · Chapter 5, Judges, Humans, and Disagreement
Disagreement, Diversity, and Topical Entropy
Human and AI disagreement is not always a defect. Sometimes it is a signal that different users value different good answers.
LLM judgeinter-rater agreementtopical entropyrubricdisagreement diversity topical entropy
What to do
- Do not fill the top results with ten pages that satisfy only one interpretation.
- Separate label noise, rubric ambiguity, judge failure, expert uncertainty, and legitimate preference diversity.
- Use cluster analysis, preference labels, slice reporting, and pairwise preference data when a single score hides meaningful groups.
Evidence to preserve
- Include one or two results for each major intent cluster so more users see something they like in the top few results.
- Preserve the inputs, versions, configurations, raw outcomes, and results for LLM judge, inter-rater agreement, topical entropy, rubric needed to reproduce work on Disagreement, Diversity, and Topical Entropy.
- Report results for LLM judge, inter-rater agreement, topical entropy, rubric by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Model disagreement should be represented explicitly. Separate label noise, rubric ambiguity, judge failure, expert uncertainty, and legitimate preference diversity. Use cluster analysis, preference labels, slice reporting, and pairwise preference data when a single score hides meaningful groups.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Disagreement, Diversity, and Topical Entropy." Testing AI Knowledge Edition, section 31.
https://jarbon.ai/testing-ai/knowledge/ch031-disagreement-diversity-topical-entropy.html