Section 172 · Chapter 20, The Practical Playbook

Minimum Viable AI Quality System

If the book feels large, start here: a small quality system that produces real evidence instead of ritual.

benchmarkdeceptionminimum viable quality

What to do

  1. Start with roughly fifty production-shaped eval cases.
  2. Use a 0-10 or 0-1 score, but define what the numbers mean.
  3. Use one LLM judge for scale, but calibrate it against human review.
  4. Review disagreements instead of hiding them.
  5. Log traces for every run: prompt, model, system message, retrieval context, tool calls, tool arguments, response, cost, latency, judge version, rubric version, and final decision.

Evidence to preserve

  • Include common use, high-value business flows, edge cases, policy boundaries, security-sensitive cases, confusing user inputs, and a few known failures from production or dogfooding.
  • Include hard blockers for privacy leaks, unsafe tool calls, policy violations, severe hallucinations, irreversible actions, and failures that would embarrass the company if screenshotted.
  • Log traces for every run: prompt, model, system message, retrieval context, tool calls, tool arguments, response, cost, latency, judge version, rubric version, and final decision.
  • Preserve the inputs, versions, configurations, raw outcomes, and results for benchmark, deception, minimum viable quality needed to reproduce work on Minimum Viable AI Quality System.

Expert note

In a real release review, treat the minimum system as an evolving control system. Version the cases, rubric, judge, model, prompts, policies, retrieval index, tools, and release thresholds together. A score without provenance is not evidence.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Minimum Viable AI Quality System." Testing AI Knowledge Edition, section 172.

https://jarbon.ai/testing-ai/knowledge/ch172-minimum-viable-quality.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400