Section 010 · Chapter 2, From Tests to Release Evidence

Stratified Reporting

Overall averages can hide weak segments. Break results down by the categories that matter.

stratified reportingRAG

What to do

  1. Define runnable checks that exercise stratified reporting and RAG.
  2. Set acceptable outcomes and blocker failures for stratified reporting and RAG before running the evaluation.
  3. Run representative cases for stratified reporting and RAG and preserve the failures that would change the decision.

Evidence to preserve

  • Report the case by slice: - legal or policy-sensitive query - local jurisdiction - renter intent - mobile result page - freshness-sensitive law - official-source requirement The eval should not only say TunedSearch scored 8.1 overall.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for stratified reporting, RAG needed to reproduce work on Stratified Reporting.
  • Report results for stratified reporting, RAG by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

When the system matters, define strata before the evaluation and ensure each important stratum has enough samples to support a decision. Too many tiny categories create noisy numbers; too few categories hide actionable risk.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Stratified Reporting." Testing AI Knowledge Edition, section 10.

https://jarbon.ai/testing-ai/knowledge/ch010-stratified-reporting.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400