Chapter 14

Frontier Safety and Containment

Treat frontier safety as a separate quality class from ordinary product bugs. Design dangerous-capability tests that measure misuse potential without teaching the dangerous content. Test for evaluation awareness, sandbagging, deception, and tool misuse with independent…

Apply this chapter

  • Treat frontier safety as a separate quality class from ordinary product bugs.
  • Design dangerous-capability tests that measure misuse potential without teaching the dangerous content.
  • Test for evaluation awareness, sandbagging, deception, and tool misuse with independent reviewers.
  • Assume containment has to hold across channels, time, operators, tools, and unknown side paths.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

7 focused briefs

Concepts in this chapter

  1. 106
    Testing Whether AI Is DangerousDo not ask vaguely whether an AI is dangerous. Test concrete hazardous capabilities, harmful behaviors, jailbreak robustness, autonomy, and…
  2. 107
    Testing CBRN and Hazardous Capability SafetyDangerous-capability testing should measure concrete misuse potential without teaching the dangerous content itself.
  3. 108
    Containment, Sandboxes, and Capability ControlIf an AI system can act, safety depends on what it is allowed to touch.
  4. 109
    Testing Manipulation, Persuasion, and Undue InfluenceA helpful assistant can become unsafe when it learns how to steer people too well.
  5. 110
    Testing Deception, Scheming, and Evaluation AwarenessThe hardest failures are not wrong answers. They are systems that behave well while watched and differently when it matters.
  6. 111
    The Gorilla Problem: Superintelligence, Containment, and UnderstandingIf a system becomes much smarter than us, containment and inspection cannot be the whole plan. The gorilla cannot audit the zookeeper.
  7. 112
    Testing Containment SystemsEvery useful AI containment system has a paradox at the center: if the system is valuable, someone or something has to interact with it. Every…

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400