Section 037 · Chapter 6, Building Evals That Matter

Real-World Evals

Public evals are useful because they make measurement concrete. They are also limited because every eval measures a particular shape of task.

real world evals

What to do

  1. Score the meta-capability, not only the chore.
  2. Score the patch, but also score whether BugPilot found the dependencies, asked for missing access, preserved payment idempotency, and left evidence another engineer could review.
  3. Define runnable checks that exercise real world evals.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for real world evals needed to reproduce work on Real-World Evals.
  • Report results for real world evals by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Real-World Evals." Testing AI Knowledge Edition, section 37.

https://jarbon.ai/testing-ai/knowledge/ch037-real-world-evals.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400