Variance, sampling, and uncertainty
Replace one-run demos with repeated measurements, slices, confidence intervals, rare-failure hunting, and release evidence.
A new book by Jason Arbon
Engineering Confidence in Non-Deterministic Systems
The practical operating manual for people who have to decide whether AI-generated code, agents, models, and products are trustworthy enough to ship.
First edition. Available now on Amazon in print and Kindle editions.
The argument
AI can produce more code, answers, designs, and decisions than any team can review the old way.
Most AI spending today still lives in experiments. The gap between an impressive demo and a production system is not usually another prompt trick. It is confidence: evidence that the system is useful, safe, economical, observable, reversible, and reliable enough for the people who depend on it.
Testing AI starts with testing because that is the doorway software teams recognize. It ends at confidence engineering: the operating discipline for shipping systems that vary from run to run, learn from changing data, call tools, personalize themselves, and behave differently under production load.
The book moves from first principles to operating practice: variance, sampling, statistical evidence, human and LLM judges, eval design, RAG, agents, generated code, security, observability, bias, interpretability, robotics, governance, and the future of AI testing itself.
Inside the book
Not a catalog of tools. Not a stack of prompt recipes. The book connects measurement, engineering, operations, security, and judgment into one practical discipline.
Replace one-run demos with repeated measurements, slices, confidence intervals, rare-failure hunting, and release evidence.
Use p-values, effect sizes, power, paired tests, calibration, abstention, and practical significance without statistical theater.
Design rubrics, calibrate humans and LLM judges, measure disagreement, manage labels, and build evals tied to product risk.
Test retrieval separately from generation, score tool trajectories, verify integration behavior, and catch code that looks right but is wrong.
Threat-model prompt injection, MCP permissions, poisoning, guardrails, dangerous capabilities, deception, and containment boundaries.
Build observability, SLOs, canaries, incident response, cost controls, robotics simulation, governance, and continuous AI-led testing.
Every chapter is also available as a short, mobile-first PDF carousel. Read it in your browser or download it to share on LinkedIn.
Companion skills
The public skills are now intentionally small: one overall Testing AI skill, plus one actionable skill for each narrative chapter. Each chapter skill has distinct trigger vocabulary and tells an AI coding agent what evidence, tests, traces, and release artifacts to produce.
Use the overall skill when you want the full confidence-engineering workflow. Use a chapter skill when the work is specific: sampling, statistical tests, judges, RAG, generated code, security, robotics, governance, or the future of AI testing AI.
Who it is for
AI is collapsing the old walls between building, testing, product judgment, and operating software. The readers differ; the release decision is shared.
Learn how to validate code and systems you did not fully write, including hidden integration, security, data, and operational failures.
Move beyond brittle exact assertions into evals, probabilistic evidence, production traces, risk slices, and human-machine review systems.
Connect prompts, models, data, retrieval, tools, policies, user experience, cost, and release controls into one testable system.
Ask for evidence that supports ship, hold, canary, rollback, investment, governance, and customer-risk decisions.
What changes after reading
You start asking the sharper questions that turn ambiguous behavior into engineering evidence.
Draft reviewers
This book improved because people took time to read rough chapters, mark confusing passages, suggest missing topics, and push the material toward practical usefulness for the broader AI quality community.
Missing topics, definitions, chatbot quality criteria, invisible dependency failures, editorial clarity, and practical reviewer perspective.
Regulatory stakes, ISO/IEC 42001 context, beginner-friendly SKILL.md explanation, reader navigation, sampling definitions, p-value clarity, and release-gate examples.
Non-deterministic framing, varying degrees of correctness and failure, and clearer calibration language.
Editorial clarity around confidence-interval language and a more beginner-friendly explanation of how intervals are calculated.
Quality dimensions, sample-size nuance, rare-failure hunting, AI judge bias questions, and practical evaluator/tooling suggestions.
Front-matter positioning, preface structure, reader framing, and a clearer bridge from traditional software quality into AI quality.
Ask AI link behavior, dense-sentence cleanup, link hygiene, repeated-example spotting, and draft-reader experience.
The suggestion to treat performance engineering as a named AI quality discipline rather than a narrow testing activity.
LLM-as-a-judge cost realism, practical eval tooling, and multi-judge voting nuance for high-stakes evaluation.
Senior-engineering feedback on tone, accountable AI builders, pass/fail nuance, judge confidence wording, and LLM-judge consistency.
Terminology consistency around testers, quality engineers, and Confidence Engineers.
Available now
Testing AI is available on Amazon in print and Kindle editions.
Buy on Amazon Amazon