Section 041 · Chapter 6, Building Evals That Matter
Stop Chasing High-Water Marks
If you rerun a noisy evaluation enough times, variance will eventually hand you a beautiful score. That does not make the system better.
variancehigh-water markRAGstop chasing high water marks
What to do
- Track every run, predefine stopping rules, preserve holdout sets, and estimate performance from the full run distribution rather than the maximum observed score.
- Define runnable checks that exercise variance, high-water mark, and RAG.
- Set acceptable outcomes and blocker failures for variance, high-water mark, and RAG before running the evaluation.
Evidence to preserve
- Track every run, predefine stopping rules, preserve holdout sets, and estimate performance from the full run distribution rather than the maximum observed score.
- Preserve the inputs, versions, configurations, raw outcomes, and results for variance, high-water mark, RAG, stop chasing high water marks needed to reproduce work on Stop Chasing High-Water Marks.
- Report results for variance, high-water mark, RAG, stop chasing high water marks by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
At scale, treat repeated evaluation as a multiple-comparisons problem. Track every run, predefine stopping rules, preserve holdout sets, and estimate performance from the full run distribution rather than the maximum observed score.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Stop Chasing High-Water Marks." Testing AI Knowledge Edition, section 41.
https://jarbon.ai/testing-ai/knowledge/ch041-stop-chasing-high-water-marks.html