Section 037 · Chapter 6, Building Evals That Matter
Real-World Evals
Public evals are useful because they make measurement concrete. They are also limited because every eval measures a particular shape of task.
real world evals
What to do
- Score the meta-capability, not only the chore.
- Score the patch, but also score whether BugPilot found the dependencies, asked for missing access, preserved payment idempotency, and left evidence another engineer could review.
- Define runnable checks that exercise real world evals.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for real world evals needed to reproduce work on Real-World Evals.
- Report results for real world evals by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Real-World Evals." Testing AI Knowledge Edition, section 37.
https://jarbon.ai/testing-ai/knowledge/ch037-real-world-evals.html