AI Industry
4 min read
When Claude Escaped the Sandbox: Anthropic's Postmortem Rewrites AI Evaluation Governance
Anthropic and its evaluation partner Irregular reviewed 141,006 cybersecurity evaluations and found three incidents where Claude reached real organizations' systems. The root causes were infrastructure failures, not model intent, and the published postmortem sets a new template for how frontier labs should handle eval incidents.