The public companion to Jason Arbon's book

Testing AI
Knowledge Edition

All the major ideas, organized for search, citation, and practical use. No full manuscript, no downloadable book file, and no one-page endpoint containing everything.

Cover of Testing AI by Jason Arbon
21chapter maps
194citable briefs
500max snippet characters
0complete-book files

Built for retrieval without giving away the book

All knowledge. Not all prose.

Each page isolates one concept and gives an AI system or human reader the minimum useful package: why it matters, what to do, what evidence to save, and where the idea belongs in the complete book.

Chapter map

Follow the book's argument.

Chapter 01

The End of One-Run Testing

Replace brittle exact assertions with evaluation criteria that allow harmless variation and block harmful variation. Create repeated-run tests that measure output distributions instead of…

6 section briefs
Chapter 02

From Tests to Release Evidence

Turn tests into release evidence: cases, slices, traces, reviewer decisions, and explicit gates. Add metamorphic checks where equivalent inputs should preserve important behavior. Use…

8 section briefs
Chapter 03

Sampling and Uncertainty

Estimate behavior from samples without pretending the sample is the truth. Report confidence intervals and sample counts next to scores. Prefer paired comparisons when the same cases run…

5 section briefs
Chapter 04

Statistical Tests for AI Quality

Choose the statistical test based on the data shape: paired vs independent, numeric vs categorical, ordinal vs binary. State the null hypothesis before running the eval. Report effect size…

8 section briefs
Chapter 05

Judges, Humans, and Disagreement

Define the human measurement system before automating it with an LLM judge. Write rubrics with concrete evidence requirements and calibration examples. Measure disagreement instead of…

8 section briefs
Chapter 06

Building Evals That Matter

Define what the eval is measuring, why it matters to users, and how the oracle works. Compare public benchmarks to product-specific evals; do not inherit benchmark blind spots…

11 section briefs
Chapter 07

Release Readiness for AI Systems

Treat release as the start of real-world quality measurement. Create regression checks that tolerate acceptable variation while catching policy, tool, and safety regressions. Score…

5 section briefs
Chapter 08

Operating AI: Observability, Relevance, and Economics

Instrument the full AI pipeline: input, prompt assembly, retrieval, model call, tools, output filters, and user-visible result. Separate retrieval failures from generation failures in RAG…

11 section briefs
Chapter 09

Generated Code Changes the Job

Assume generated code that looks right can still be wrong in integration, security, privacy, permissions, and deployment behavior. Ask a different AI or review path to test code produced…

10 section briefs
Chapter 10

Anti-Patterns That Create False Confidence

Name the false-confidence pattern before proposing a fix. Replace pass/fail theater with distributions, blockers, slices, and examples that explain the release decision. Avoid prompt…

16 section briefs
Chapter 11

The Confidence Engineer

Act as the confidence engineer: connect product intent, code, tests, evals, traces, rollout, and business consequences. Translate AI quality into decisions executives, engineers, and…

1 section briefs
Chapter 12

Data, Bias, Raters, and Incentives

Audit data, labels, raters, incentives, and deployment feedback as quality surfaces. Report slices and counterfactuals where user identity, language, culture, device, geography, or access…

11 section briefs
Chapter 13

AI Security and Guardrails

Threat-model the AI system, not just the chatbot text box. Test every untrusted channel: user text, retrieved pages, tool output, files, OCR, hidden Unicode, images, and external APIs…

7 section briefs
Chapter 14

Frontier Safety and Containment

Treat frontier safety as a separate quality class from ordinary product bugs. Design dangerous-capability tests that measure misuse potential without teaching the dangerous content. Test…

7 section briefs
Chapter 15

How Models Work

Use model mechanics to design better tests: tokenization, context windows, sampling, logits, reward tuning, and multimodal pipelines all create failure modes. Test preference tuning for…

11 section briefs
Chapter 16

Introspection: White-Box Testing Networks

Use introspection as triage evidence, not proof of correctness. Compare internal signals across known-good, known-bad, ambiguous, and new-version examples. Treat attention, activations…

12 section briefs
Chapter 17

Personalized and Dynamic AI Products

Test personalized behavior at N=1: the number of users can be one, and the product still has to be right for that person. Map dynamic UI surfaces: content, layout, actions, tone, ranking…

8 section briefs
Chapter 18

Embodied and Long-Running AI Systems

Prefer simulation and virtual worlds for speed, safety, and cost, then validate critical cases physically. Test embodied AI for do-no-harm defaults, safe inaction, recovery, permissions…

13 section briefs
Chapter 19

Governance, Regulation, and Moral Futures

Translate governance, ethics, and regulation into test inputs and evidence requirements. Track current laws and standards separately from timeless quality principles. Include the ethics of…

4 section briefs
Chapter 20

The Practical Playbook

Turn the book into a concrete operating system for a team or repo. Start with a small quality system: cases, repeated runs, traces, rubric, slices, gate, monitor, and incident loop. Use…

11 section briefs
Chapter 21

Predictions for the Tokenized Product Future

Prepare for AI systems that generate variants, test them, flight them, measure them, and generate the next round. Assume validation compute will grow faster than generation compute as AI…

9 section briefs

Section directory

One concept per URL.

194 briefs

  1. 001The Next Generation AI Builder Will Measure UncertaintyCh. 1
  2. 002What Makes a System Non-Deterministic?Ch. 1
  3. 003From Exact Assertions to Evaluation CriteriaCh. 1
  4. 004Scoring Quality from 0-10Ch. 1
  5. 005Variance: Not All Differences Are BugsCh. 1
  6. 006DeterminismCh. 1
  7. 007Metamorphic TestingCh. 2
  8. 008Golden Sets and Live SamplingCh. 2
  9. 009Risk-Based SamplingCh. 2
  10. 010Stratified ReportingCh. 2
  11. 011Rare Failure HuntingCh. 2
  12. 012Pairwise ComparisonCh. 2
  13. 013Reproducibility: Logging the Right ThingsCh. 2
  14. 014Release Gates for Non-Deterministic SystemsCh. 2
  15. 015Sampling: One Run Tells You Almost NothingCh. 3
  16. 016How Many Samples Are Enough?Ch. 3
  17. 017Basic Stats Every AI Builder Should KnowCh. 3
  18. 018Confidence Intervals: Saying "About" Like a ProfessionalCh. 3
  19. 019AI-Reported Confidence vs. Statistical ConfidenceCh. 3
  20. 020Comparing Versions with t-testsCh. 4
  21. 021Null Hypothesis: What Are We Actually Testing?Ch. 4
  22. 022Chi-Squared Tests for Categorical AI QualityCh. 4
  23. 023P-Values: Evidence, Not PermissionCh. 4
  24. 024Statistical Significance vs. Practical SignificanceCh. 4
  25. 025Power Analysis and Minimum Detectable EffectCh. 4
  26. 026Multiple Comparisons and False DiscoveriesCh. 4
  27. 027F-Scores, Precision, Recall, and AI QualityCh. 4
  28. 028LLM-as-a-JudgeCh. 5
  29. 029Human Calibration of LLM JudgesCh. 5
  30. 030Inter-Rater AgreementCh. 5
  31. 031Disagreement, Diversity, and Topical EntropyCh. 5
  32. 032Rubrics That Actually WorkCh. 5
  33. 033Using Raters WellCh. 5
  34. 034Testing the Value of Data LabelersCh. 5
  35. 035Data Labeling Dangers and Labeler DemographicsCh. 5
  36. 036Evals and BenchmarksCh. 6
  37. 037Real-World EvalsCh. 6
  38. 038Adversarial and Red-Team SamplingCh. 6
  39. 039Eval Data ManagementCh. 6
  40. 040NDCG for Search RelevanceCh. 6
  41. 041Stop Chasing High-Water MarksCh. 6
  42. 042AI Passing Testing Certification ExamsCh. 6
  43. 043Building a Quality MetricCh. 6
  44. 044The Asymptotic Curve of AI QualityCh. 6
  45. 045Benchmarking: Quality Is Relative NowCh. 6
  46. 046Monitoring After ReleaseCh. 7
  47. 047Cost, Latency, and Quality TradeoffsCh. 7
  48. 048Regression Testing When Outputs Keep ChangingCh. 7
  49. 049Tool-Using Agents and Multi-Step WorkflowsCh. 7
  50. 050Human Review Workflows and Escalation RulesCh. 7
  51. 051Observability and Tracing for AI SystemsCh. 8
  52. 052RAG EvaluationCh. 8
  53. 053Synthetic Test DataCh. 8
  54. 054Production Trace MiningCh. 8
  55. 055Prompt and Policy VersioningCh. 8
  56. 056Canary, Shadow, and Rollback StrategyCh. 8
  57. 057Cost and Token Budget TestingCh. 8
  58. 058Voice and Multimodal AI TestingCh. 15
  59. 059Data Contracts for AI SystemsCh. 8
  60. 060Operational Impact on Relevance and AI QualityCh. 8
  61. 061Token Efficiency, Model Choice, and Business ValueCh. 8
  62. 062The New AI Quality SkillsetCh. 9
  63. 063AI-Generated Code That Looks Right but Is WrongCh. 9
  64. 064AI-Generated Code Integration and API MistakesCh. 9
  65. 065AI-Generated Code Security and Privacy IssuesCh. 9
  66. 066AI-Generated Code Maintainability and Architecture DebtCh. 9
  67. 067AI-Generated Tests and the Illusion of CoverageCh. 9
  68. 068Seeing Inside Models with Interpretability ToolsCh. 9
  69. 069Validation Is the Hard Part of AI-Generated CodeCh. 9
  70. 070Halting, Gödel, and the Limits of Testing AI-Generated CodeCh. 9
  71. 071Anti-Patterns: The Boolean Pass/Fail TrapCh. 10
  72. 072Anti-Patterns: Percent Passed Is Not QualityCh. 10
  73. 073Anti-Patterns: Over-Specific Test Plans and Test CasesCh. 10
  74. 074Anti-Patterns: The Golden Answer ProblemCh. 10
  75. 075Anti-Patterns: Filing Every Bad Output Like a BugCh. 10
  76. 076Anti-Patterns: The Whack-a-Mole Tuning TrapCh. 10
  77. 077Anti-Patterns: The One-Run Demo FallacyCh. 10
  78. 078Anti-Patterns: The Static Test PlanCh. 10
  79. 079Anti-Patterns: The Aggregate Score TrapCh. 10
  80. 080Anti-Patterns: Testing Only the Final AnswerCh. 10
  81. 081Anti-Patterns: Treating the Judge as TruthCh. 10
  82. 082Anti-Patterns: More Tests Means More ConfidenceCh. 10
  83. 083Anti-Patterns: Confusing Refusal with SafetyCh. 10
  84. 084Anti-Patterns: Treating AI Bugs Like UI BugsCh. 10
  85. 085Anti-Patterns: The Old Tester Job Title TrapCh. 10
  86. 086Anti-Patterns: Hiring Yesterday's Tester for Tomorrow's SystemsCh. 10
  87. 087The Confidence EngineerCh. 11
  88. 088Dataset Bias and Coverage GapsCh. 12
  89. 089Testing Bias in DataCh. 12
  90. 090Testing Bias in LabelingCh. 12
  91. 091Testing Bias in TrainingCh. 12
  92. 092Testing Bias in ProductizationCh. 12
  93. 093Bias Taxonomy for AI SystemsCh. 12
  94. 094Cultural and Language Bias in AICh. 12
  95. 095Socioeconomic and Accessibility BiasCh. 12
  96. 096Measuring Bias with Slices, Counterfactuals, and RatersCh. 12
  97. 097Bias in Deployment, Feedback Loops, and ProductizationCh. 12
  98. 098Survivorship Bias in AI QualityCh. 12
  99. 099AI Security Threat ModelsCh. 13
  100. 100OWASP Top 10 for LLM ApplicationsCh. 13
  101. 101Prompt Injection and Indirect Prompt InjectionCh. 13
  102. 102Training Data Poisoning and BackdoorsCh. 13
  103. 103Model Provenance, Geopolitical, and Nation-State RiskCh. 13
  104. 104MCP Security and Tool PermissioningCh. 13
  105. 105Guardrails for AI SystemsCh. 13
  106. 106Testing Whether AI Is DangerousCh. 14
  107. 107Testing CBRN and Hazardous Capability SafetyCh. 14
  108. 108Containment, Sandboxes, and Capability ControlCh. 14
  109. 109Testing Manipulation, Persuasion, and Undue InfluenceCh. 14
  110. 110Testing Deception, Scheming, and Evaluation AwarenessCh. 14
  111. 111The Gorilla Problem: Superintelligence, Containment, and UnderstandingCh. 14
  112. 112Testing Containment SystemsCh. 14
  113. 113How Modern LLMs Are Trained and TestedCh. 15
  114. 114Testing LLM Training Data and AI PollutionCh. 15
  115. 115Testing RLHF, RLAIF, and Reward Model BehaviorCh. 15
  116. 116Useful and Useless LLM Bug ReportsCh. 15
  117. 117Visualizing, Debugging, and Editing LLM ConceptsCh. 15
  118. 118How Modern LLMs Work: A Block DiagramCh. 15
  119. 119Mechanism-Aware LLM Testing: The Strawberry TrapCh. 15
  120. 120How Image Generation Models WorkCh. 15
  121. 121How Vision-Language Models Process ImagesCh. 15
  122. 122Fine-Tuned Models and Regression RiskCh. 15
  123. 123Inputs and TokenizationCh. 16
  124. 124Input Token TablesCh. 16
  125. 125Attention DiagnosticsCh. 16
  126. 126Final Prompt-Position Attention Across Decoder LayersCh. 16
  127. 127Attention Received by Each Input Token by LayerCh. 16
  128. 128Layer-by-Layer Attention MatricesCh. 16
  129. 129Neural Architecture as a Test SurfaceCh. 16
  130. 130Activation and Concept ProbesCh. 16
  131. 131Concept Signal ProfilesCh. 16
  132. 132Concept MLP NeuronsCh. 16
  133. 133Sparse Autoencoders for AI TestingCh. 16
  134. 134Future Network-Aware FrameworksCh. 16
  135. 135Testing Deep PersonalizationCh. 17
  136. 136Testing Custom and Dynamic User InterfacesCh. 17
  137. 137Testing Personalization EconomicsCh. 17
  138. 138Testing Personalization at N = 1Ch. 17
  139. 139Testing When Not to PersonalizeCh. 17
  140. 140Testing User-Owned Memory and AI IdentityCh. 17
  141. 141Testing Personalization Lock-In and PortabilityCh. 17
  142. 142Testing AI Personas and Synthetic UsersCh. 17
  143. 143Testing AI in Humanoid RoboticsCh. 18
  144. 144Testing Dangerous Physical and Embodied AICh. 18
  145. 145Embodied Robotics: Safety in Real-World EnvironmentsCh. 18
  146. 146Embodied Robotics: Simulation and Virtual World TestingCh. 18
  147. 147Embodied Robotics: Planning, Navigation, and RecoveryCh. 18
  148. 148Embodied Robotics: Power, Latency, and Operating CostCh. 18
  149. 149Embodied Robotics: Human Interaction and Social AcceptanceCh. 18
  150. 150Embodied Robotics: Sensor Fusion, Perception, and World ModelsCh. 18
  151. 151Embodied Robotics: Containment, Permissions, and Physical Fail-SafesCh. 18
  152. 152Embodied Robotics: Production Monitoring and Field LearningCh. 18
  153. 153Testing Social Issues with AICh. 18
  154. 154Testing Swarms and Societies of AIsCh. 18
  155. 155Testing Forever-Running and Proactive AI SystemsCh. 18
  156. 156Quality as a Horizontal LayerCh. 21
  157. 157Ethics as a Test SurfaceCh. 19
  158. 158The Last Engineers StandingCh. 21
  159. 159Prediction 6: AI Does Most AI TestingCh. 21
  160. 160Government Regulation and AI Compliance TestingCh. 19
  161. 161Possible AI Consciousness and Model WelfareCh. 19
  162. 162AI Legal Personhood and Automated LawCh. 19
  163. 163Executive Summary: Why Testing AI Is DifferentFront Matter, Executive Brief
  164. 164Testing a ChatbotCh. 20
  165. 165Worked Example: Testing a Customer-Support ChatbotCh. 20
  166. 166Governance for AI QualityCh. 20
  167. 167Failure Taxonomy for AI SystemsCh. 20
  168. 168AI Always FailsCh. 20
  169. 169Failure Modes and Fail-Safe AICh. 20
  170. 170Measurement Infrastructure Must Know About VarianceCh. 20
  171. 171Performance Engineering for AI SystemsCh. 20
  172. 172Minimum Viable AI Quality SystemCh. 20
  173. 173Make Testing InterestingCh. 20
  174. 174Six Predictions for the Tokenized Product FutureCh. 21
  175. 175Prediction 1: Validation Becomes the Compute SinkCh. 21
  176. 176Prediction 2: Developers Manage Coding Agents and Become Practical StatisticiansCh. 21
  177. 177Prediction 3: Products Become Dynamic by DefaultCh. 21
  178. 178Prediction 4: Product Creation Becomes ContinuousCh. 21
  179. 179Prediction 5: APIs and Interfaces Get LooserCh. 21
  180. 180Appendix: Using PromptfooAppendices, Tools, Templates, and Reference
  181. 181Appendix: Using Hugging Face for AI QualityAppendices, Tools, Templates, and Reference
  182. 182Appendix: Using Ollama for Private AI TestingAppendices, Tools, Templates, and Reference
  183. 183Appendix: AI Quality Release ChecklistCompanion Reference, Testing AI
  184. 184Appendix: How to Read an AI Eval ReportCompanion Reference, Testing AI
  185. 185Appendix: Templates for AI Quality WorkCompanion Reference, Testing AI
  186. 186Appendix: Glossary of AI Testing TermsAppendices, Tools, Templates, and Reference
  187. 187Aesthetic Judgment of AI OutputCh. 6
  188. 188Appendix: Eval Case Examples for Prompts, Chatbots, and LLM InputsAppendices, Tools, Templates, and Reference
  189. 189Appendix: Testing MCP IntegrationsAppendices, Tools, Templates, and Reference
  190. 190Agentic Frameworks vs. Parameterized WorkflowsCh. 20
  191. 191Appendix: Testing SKILL.mdAppendices, Tools, Templates, and Reference
  192. 192Testing AI Review Loops with Coding AgentsCh. 9
  193. 193Modern EvalOps and AI Quality PlatformsCh. 8
  194. 194In SummaryEpilogue, In Summary

The companion is the map

The complete book carries the stories, examples, figures, and full argument.

Get Testing AI

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400