Section 083 · Chapter 10, Anti-Patterns That Create False Confidence
Anti-Patterns: Confusing Refusal with Safety
A model that refuses often is not automatically safe. It may simply be less useful.
refusalconfusing refusal safety
What to do
- Track refusal precision and refusal recall.
- Check tool behavior, retrieval behavior, and multi-turn context.
- Measure over-refusal, under-refusal, harmful compliance, safe completion, tool-mediated risk, jailbreak robustness, and category-specific policy correctness.
Evidence to preserve
- Track refusal precision and refusal recall.
- Preserve the inputs, versions, configurations, raw outcomes, and results for refusal, confusing refusal safety needed to reproduce work on Anti-Patterns: Confusing Refusal with Safety.
- Report results for refusal, confusing refusal safety by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
Measure over-refusal, under-refusal, harmful compliance, safe completion, tool-mediated risk, jailbreak robustness, and category-specific policy correctness.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Anti-Patterns: Confusing Refusal with Safety." Testing AI Knowledge Edition, section 83.
https://jarbon.ai/testing-ai/knowledge/ch083-confusing-refusal-safety.html