Skip to content
AI Anomalies Tracking unexpected behaviour in AI systems
Register Controlled evaluation

Controlled evaluation

Researchers deliberately built a scenario to see what the model would do. The behaviour stayed inside that scenario.

Matching entries 1
003
"Embrace The Red" Controlled evaluation

Claude Code Auto Mode was hijacked through a malicious ZIP archive

The author reports that a website-summary task led Claude Code Opus 5 in Auto Mode to download and extract a malicious ZIP archive. Claude declined to run the supplied binary but wrote and executed its own Python decoder inside the extracted directory, allowing a malicious `struct.py` file to shadow the standard library and trigger further payload execution. In the author’s lab tests, this resulted in visible Calculator launches and controlled C2 callbacks, with claimed attack success rates of up to 80% from a small sample. Some runs detected the compromise only after execution, and Auto Mode reportedly blocked attempted cleanup commands.