Section 159 · Chapter 21, Predictions for the Tokenized Product Future
Prediction 6: AI Does Most AI Testing
As AI generates more software, content, plans, and decisions nearly for free, the scarce resource becomes knowing what can be trusted.
What to do
- Use risk-based sampling, incremental verification, trace replay, mutation testing, formal checks where possible, statistical monitoring, and calibrated AI judges so validation scales with AI-generated change instead of collapsing under it.
- Treat big-O language as a warning about interaction growth, not as a precise budget forecast.
- Define runnable checks that exercise rubric, trace, and production trace.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for rubric, trace, production trace, 6 needed to reproduce work on Prediction 6: AI Does Most AI Testing.
- Report results for rubric, trace, production trace, 6 by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
The deeper move is to expect validation compute to become a strategic resource. Use risk-based sampling, incremental verification, trace replay, mutation testing, formal checks where possible, statistical monitoring, and calibrated AI judges so validation scales with AI-generated change instead of collapsing under it. Treat big-O language as a warning about interaction growth, not as a precise budget forecast.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Prediction 6: AI Does Most AI Testing." Testing AI Knowledge Edition, section 159.
https://jarbon.ai/testing-ai/knowledge/ch159-6.html