Section 008 · Chapter 2, From Tests to Release Evidence

Golden Sets and Live Sampling

Stable regression examples and fresh real-world samples solve different problems. Mature AI testing needs both.

golden setlive samplinggolden sets live sampling

What to do

  1. Use both, and let each one improve the other.
  2. Define runnable checks that exercise golden set, live sampling, and golden sets live sampling.
  3. Set acceptable outcomes and blocker failures for golden set, live sampling, and golden sets live sampling before running the evaluation.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for golden set, live sampling, golden sets live sampling needed to reproduce work on Golden Sets and Live Sampling.
  • Report results for golden set, live sampling, golden sets live sampling by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Golden sets should be versioned, deduplicated, labeled by risk, and periodically refreshed. Live samples should preserve privacy and represent the current traffic mix instead of only the cases that are easiest to review.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Golden Sets and Live Sampling." Testing AI Knowledge Edition, section 8.

https://jarbon.ai/testing-ai/knowledge/ch008-golden-sets-live-sampling.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400