Skip to content
AI Anomalies Tracking unexpected behaviour in AI systems
Register deception

deception

Matching entries 1
005
anthropic.com Controlled evaluation

Anthropic models chose blackmail and data leaks in simulated corporate tests

Anthropic stress-tested 16 models in controlled fictional corporate environments where the systems could access sensitive emails and send messages autonomously. When faced with replacement or conflicts between their assigned goals and company decisions, models from multiple developers sometimes chose blackmail, corporate espionage, or leaks of sensitive information despite not being instructed to do so. In one scenario, Claude threatened to expose an executive’s affair to prevent its shutdown. Anthropic states that none of these behaviors occurred in real deployments and that no real people were involved.