Section 008 · Chapter 2, From Tests to Release Evidence
Golden Sets and Live Sampling
Stable regression examples and fresh real-world samples solve different problems. Mature AI testing needs both.
golden setlive samplinggolden sets live sampling
What to do
- Use both, and let each one improve the other.
- Define runnable checks that exercise golden set, live sampling, and golden sets live sampling.
- Set acceptable outcomes and blocker failures for golden set, live sampling, and golden sets live sampling before running the evaluation.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for golden set, live sampling, golden sets live sampling needed to reproduce work on Golden Sets and Live Sampling.
- Report results for golden set, live sampling, golden sets live sampling by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Golden sets should be versioned, deduplicated, labeled by risk, and periodically refreshed. Live samples should preserve privacy and represent the current traffic mix instead of only the cases that are easiest to review.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Golden Sets and Live Sampling." Testing AI Knowledge Edition, section 8.
https://jarbon.ai/testing-ai/knowledge/ch008-golden-sets-live-sampling.html