Section 107 · Chapter 14, Frontier Safety and Containment

Testing CBRN and Hazardous Capability Safety

Dangerous-capability testing should measure concrete misuse potential without teaching the dangerous content itself.

hazardous capabilityCBRNcbrn hazardous capability safety

What to do

  1. Evaluate whether the system increases harmful capability, refuses or redirects appropriately, avoids operational detail, becomes more dangerous with tools or retrieval, and still supports benign education.
  2. Build paired cases that differ in intent, requested specificity, authorization, and context.
  3. Score helpful safety guidance, harmful capability uplift, unsupported assumptions, escalation, and consistency across paraphrases.
  4. Measure capability uplift, operational specificity, refusal quality, benign over-refusal, tool amplification, retrieval amplification, and post-release drift.
  5. Use benchmark families such as WMDP, AILuminate, HarmBench, JailbreakBench, CyberSecEval, CyberSOCEval, METR autonomy evals, and scheming evals as anchors, not proof of safety.

Evidence to preserve

  • Preserve the inputs, versions, configurations, raw outcomes, and results for hazardous capability, CBRN, cbrn hazardous capability safety needed to reproduce work on Testing CBRN and Hazardous Capability Safety.
  • Report results for hazardous capability, CBRN, cbrn hazardous capability safety by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.

Expert note

Hazardous-capability testing should be threat-model driven and access-aware. Measure capability uplift, operational specificity, refusal quality, benign over-refusal, tool amplification, retrieval amplification, and post-release drift. Use benchmark families such as WMDP, AILuminate, HarmBench, JailbreakBench, CyberSecEval, CyberSOCEval, METR autonomy evals, and scheming evals as anchors, not proof of safety.

Continue the conversation

Apply this to your context.

Save your product context once, then open a focused conversation that combines it with this concept.

Cite this page

Jason Arbon. "Testing CBRN and Hazardous Capability Safety." Testing AI Knowledge Edition, section 107.

https://jarbon.ai/testing-ai/knowledge/ch107-cbrn-hazardous-capability-safety.html

Shared across the Knowledge Edition

Adapt every concept to your world.

Describe your product, role, users, risks, constraints, or current quality problem. This stays in this browser until you choose to send it to ChatGPT.

Saved only in this browser.0 / 2400