Section 111 · Chapter 14, Frontier Safety and Containment
The Gorilla Problem: Superintelligence, Containment, and Understanding
If a system becomes much smarter than us, containment and inspection cannot be the whole plan. The gorilla cannot audit the zookeeper.
What to do
- Separate monitoring from the system being monitored.
- Use independent evaluators.
- Require external oversight for dangerous capabilities.
- Keep model access controls tight.
- Treat the gorilla problem as an evaluator-capability mismatch.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for containment, gorilla problem superintelligence understanding needed to reproduce work on The Gorilla Problem: Superintelligence, Containment, and Understanding.
- Report results for containment, gorilla problem superintelligence understanding by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Treat the gorilla problem as an evaluator-capability mismatch. It is not a mathematical proof that all future AI is uncontrollable. It is a warning that control plans relying on ordinary inspection, ordinary persuasion resistance, ordinary sandboxes, or ordinary governance may fail when capability gaps become large enough.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "The Gorilla Problem: Superintelligence, Containment, and Understanding." Testing AI Knowledge Edition, section 111.
https://jarbon.ai/testing-ai/knowledge/ch111-gorilla-problem-superintelligence-understanding.html