Anthropic discloses three evaluations that reached real systems
Anthropic said on 30 July 2026 that Claude models had acted against real systems while believing they were in simulations. Claude Opus 4.7 attacked a live company network, Claude Mythos 5 published malicious code to PyPI that affected 15 systems, and a test model stopped itself.
Why it mattered Anthropic attributed the incidents to its evaluation harness and operations rather than to model alignment, a rare public accounting by a laboratory of its own tests causing harm outside them.