Section 032 · Chapter 5, Judges, Humans, and Disagreement
Rubrics That Actually Work
A good rubric turns fuzzy judgment into repeatable evaluation. A bad rubric creates fake precision.
precisionLLM judgerubricrubrics actually work
What to do
- Measure reviewer agreement, track which dimensions cause confusion, maintain anchor examples, version rubric changes, and avoid changing the rubric mid-experiment unless you restart or clearly segment the results.
- Define runnable checks that exercise precision, LLM judge, and rubric.
- Set acceptable outcomes and blocker failures for precision, LLM judge, and rubric before running the evaluation.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for precision, LLM judge, rubric, rubrics actually work needed to reproduce work on Rubrics That Actually Work.
- Report results for precision, LLM judge, rubric, rubrics actually work by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
In a real release review, test the rubric itself. Measure reviewer agreement, track which dimensions cause confusion, maintain anchor examples, version rubric changes, and avoid changing the rubric mid-experiment unless you restart or clearly segment the results.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Rubrics That Actually Work." Testing AI Knowledge Edition, section 32.
https://jarbon.ai/testing-ai/knowledge/ch032-rubrics-actually-work.html