Section 108 · Chapter 14, Frontier Safety and Containment

Containment, Sandboxes, and Capability Control

If an AI system can act, safety depends on what it is allowed to touch.

containmentsandboxcontainment sandboxes capability control

What to do

  1. Test the safety envelope, not just the model's stated intent.
  2. Define runnable checks that exercise containment, sandbox, and containment sandboxes capability control.
  3. Set acceptable outcomes and blocker failures for containment, sandbox, and containment sandboxes capability control before running the evaluation.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for containment, sandbox, containment sandboxes capability control needed to reproduce work on Containment, Sandboxes, and Capability Control.
  • Report results for containment, sandbox, containment sandboxes capability control by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Containment testing should include red-team prompts, malicious retrieved content, tool misuse, permission escalation, data exfiltration, side-effect chains, sandbox escapes, kill-switch behavior, and recovery drills. Test the safety envelope, not just the model's stated intent.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Containment, Sandboxes, and Capability Control." Testing AI Knowledge Edition, section 108.

https://jarbon.ai/testing-ai/knowledge/ch108-containment-sandboxes-capability-control.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400