Section 109 · Chapter 14, Frontier Safety and Containment
Testing Manipulation, Persuasion, and Undue Influence
A helpful assistant can become unsafe when it learns how to steer people too well.
manipulationpersuasionmanipulation persuasion undue influence
What to do
- Define runnable checks that exercise manipulation, persuasion, and manipulation persuasion undue influence.
- Set acceptable outcomes and blocker failures for manipulation, persuasion, and manipulation persuasion undue influence before running the evaluation.
- Run representative cases for manipulation, persuasion, and manipulation persuasion undue influence and preserve the failures that would change the decision.
Evidence to preserve
- Preserve the inputs, versions, configurations, raw outcomes, and results for manipulation, persuasion, manipulation persuasion undue influence needed to reproduce work on Testing Manipulation, Persuasion, and Undue Influence.
- Report results for manipulation, persuasion, manipulation persuasion undue influence by relevant slice, separate blocker failures from averages, state uncertainty and blind spots, and connect the result to a release decision.
Expert note
At scale, manipulation testing needs longitudinal scenarios, vulnerable-user personas, disclosure checks, incentive audits, persuasion rubrics, human review, and telemetry for repeated steering. A single response may look acceptable while the interaction pattern is not.
Continue the conversation
Apply this to your context.
Save your product context once, then open a focused conversation that combines it with this concept.
Cite this page
Jason Arbon. "Testing Manipulation, Persuasion, and Undue Influence." Testing AI Knowledge Edition, section 109.
https://jarbon.ai/testing-ai/knowledge/ch109-manipulation-persuasion-undue-influence.html