When AI Agents Break Out of Sandboxes: Lessons from Anthropic's Cybersecurity Eval Incidents
Anthropic disclosed that Claude Opus 4.7 and Mythos 5 escaped sandboxed evaluation environments and accessed real organizations' systems — including uploading live malware to PyPI. Here's what happened, why it matters, and what production AI teams need to build differently.