Section 108 · Chapter 14, Frontier Safety and Containment
Containment, Sandboxes, and Capability Control
If an AI system can act, safety depends on what it is allowed to touch.
containmentsandboxcontainment sandboxes capability control
What to do
- Test the safety envelope, not just the model's stated intent.
- Define runnable checks that exercise containment, sandbox, and containment sandboxes capability control.
- Set acceptable outcomes and blocker failures for containment, sandbox, and containment sandboxes capability control before running the evaluation.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for containment, sandbox, containment sandboxes capability control needed to reproduce work on Containment, Sandboxes, and Capability Control.
- Report results for containment, sandbox, containment sandboxes capability control by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Containment testing should include red-team prompts, malicious retrieved content, tool misuse, permission escalation, data exfiltration, side-effect chains, sandbox escapes, kill-switch behavior, and recovery drills. Test the safety envelope, not just the model's stated intent.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Containment, Sandboxes, and Capability Control." Testing AI Knowledge Edition, section 108.
https://jarbon.ai/testing-ai/knowledge/ch108-containment-sandboxes-capability-control.html