Section 039 · Chapter 6, Building Evals That Matter

Eval Data Management

If prompts, datasets, rubrics, labels, judges, and model versions are not versioned, the evaluation cannot be trusted.

rubrictraceeval data management

What to do

  1. Define runnable checks that exercise rubric, trace, and eval data management.
  2. Set acceptable outcomes and blocker failures for rubric, trace, and eval data management before running the evaluation.
  3. Run representative cases for rubric, trace, and eval data management and preserve the failures that would change the decision.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for rubric, trace, eval data management needed to reproduce work on Eval Data Management.
  • Report results for rubric, trace, eval data management by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Expert teams treat evals like experiments and production telemetry at the same time. They keep immutable run records, separate raw data from derived labels, document schema changes, and make comparisons only between compatible runs.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Eval Data Management." Testing AI Knowledge Edition, section 39.

https://jarbon.ai/testing-ai/knowledge/ch039-eval-data-management.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400