Section 032 · Chapter 5, Judges, Humans, and Disagreement

Rubrics That Actually Work

A good rubric turns fuzzy judgment into repeatable evaluation. A bad rubric creates fake precision.

precisionLLM judgerubricrubrics actually work

What to do

  1. Measure reviewer agreement, track which dimensions cause confusion, maintain anchor examples, version rubric changes, and avoid changing the rubric mid-experiment unless you restart or clearly segment the results.
  2. Define runnable checks that exercise precision, LLM judge, and rubric.
  3. Set acceptable outcomes and blocker failures for precision, LLM judge, and rubric before running the evaluation.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for precision, LLM judge, rubric, rubrics actually work needed to reproduce work on Rubrics That Actually Work.
  • Report results for precision, LLM judge, rubric, rubrics actually work by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

In a real release review, test the rubric itself. Measure reviewer agreement, track which dimensions cause confusion, maintain anchor examples, version rubric changes, and avoid changing the rubric mid-experiment unless you restart or clearly segment the results.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Rubrics That Actually Work." Testing AI Knowledge Edition, section 32.

https://jarbon.ai/testing-ai/knowledge/ch032-rubrics-actually-work.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400