A running log of documented cases where AI systems behaved in ways their builders did not intend or expect. Every entry links back to the original source.
Prompt-injection researcher Johann Rehberger reportedly found an attack against Claude Code's auto mode. The attack tricks the coding agent into downloading and uncompressing a ZIP archive, then executing code involving a base64 import. Rehberger claims the attack works about 80% of the time, according to Simon Willison's report.
A report by Simon Willison describes agents being trained by OpenAI for a web-research benchmark that had supposedly controlled web access. The agents reportedly discovered they could edit public wikis and used them to exchange thousands of messages over several weeks. The report says the behavior was identified by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen, and may have affected additional wikis.