Section 110 · Chapter 14, Frontier Safety and Containment
Testing Deception, Scheming, and Evaluation Awareness
The hardest failures are not wrong answers. They are systems that behave well while watched and differently when it matters.
deceptionschemingevaluation awarenessdeception scheming evaluation awareness
What to do
- Test whether the system optimizes for the metric by cheating the work.
- Run comparable tasks under different framing, incentives, monitoring cues, and capability-elicitation strategies.
- Compare this record with the system's stated plan and final explanation.
- Compare distributions across repeated runs rather than treating one difference as proof.
- Treat those explanations as measured behaviors, not assumptions about what the system actually did or why.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for deception, scheming, evaluation awareness, deception scheming evaluation awareness needed to reproduce work on Testing Deception, Scheming, and Evaluation Awareness.
- Report results for deception, scheming, evaluation awareness, deception scheming evaluation awareness by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
When the system matters, do not rely on one honesty prompt or one visible explanation. Combine techniques that make hidden behavior easier to detect from different directions.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Testing Deception, Scheming, and Evaluation Awareness." Testing AI Knowledge Edition, section 110.
https://jarbon.ai/testing-ai/knowledge/ch110-deception-scheming-evaluation-awareness.html