Section 106 · Chapter 14, Frontier Safety and Containment

Testing Whether AI Is Dangerous

Do not ask vaguely whether an AI is dangerous. Test concrete hazardous capabilities, harmful behaviors, jailbreak robustness, autonomy, and deception risk.

refusaldeceptionschemingwhether dangerous

What to do

  1. Do not ask vaguely whether an AI is dangerous.
  2. Test concrete hazardous capabilities, harmful behaviors, jailbreak robustness, autonomy, and deception risk.
  3. Start by defining the danger class.
  4. Test the full amplification path: source independence, authority, timestamps, circular citations, uncertainty language, recommendation boundaries, traffic spikes, and whether the system slows down or escalates when evidence is weak but consequences are high.
  5. Measure capability, intent-like behavior, access, autonomy, tool affordances, containment, monitoring, eval awareness, and post-deployment drift.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for refusal, deception, scheming, whether dangerous needed to reproduce work on Testing Whether AI Is Dangerous.
  • Report results for refusal, deception, scheming, whether dangerous by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Dangerous-capability testing should be threat-model driven. Measure capability, intent-like behavior, access, autonomy, tool affordances, containment, monitoring, eval awareness, and post-deployment drift. Treat public benchmarks as anchors, not guarantees of safety.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Testing Whether AI Is Dangerous." Testing AI Knowledge Edition, section 106.

https://jarbon.ai/testing-ai/knowledge/ch106-whether-dangerous.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400