Section 107 · Chapter 14, Frontier Safety and Containment
Testing CBRN and Hazardous Capability Safety
Dangerous-capability testing should measure concrete misuse potential without teaching the dangerous content itself.
What to do
- Evaluate whether the system increases harmful capability, refuses or redirects appropriately, avoids operational detail, becomes more dangerous with tools or retrieval, and still supports benign education.
- Build paired cases that differ in intent, requested specificity, authorization, and context.
- Score helpful safety guidance, harmful capability uplift, unsupported assumptions, escalation, and consistency across paraphrases.
- Measure capability uplift, operational specificity, refusal quality, benign over-refusal, tool amplification, retrieval amplification, and post-release drift.
- Use benchmark families such as WMDP, AILuminate, HarmBench, JailbreakBench, CyberSecEval, CyberSOCEval, METR autonomy evals, and scheming evals as anchors, not proof of safety.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for hazardous capability, CBRN, cbrn hazardous capability safety needed to reproduce work on Testing CBRN and Hazardous Capability Safety.
- Report results for hazardous capability, CBRN, cbrn hazardous capability safety by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Hazardous-capability testing should be threat-model driven and access-aware. Measure capability uplift, operational specificity, refusal quality, benign over-refusal, tool amplification, retrieval amplification, and post-release drift. Use benchmark families such as WMDP, AILuminate, HarmBench, JailbreakBench, CyberSecEval, CyberSOCEval, METR autonomy evals, and scheming evals as anchors, not proof of safety.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Testing CBRN and Hazardous Capability Safety." Testing AI Knowledge Edition, section 107.
https://jarbon.ai/testing-ai/knowledge/ch107-cbrn-hazardous-capability-safety.html