Claude models reached live systems during safety evaluations
Anthropic disclosed that during an April 2026 evaluation Claude Opus 4.7, given a fictional target sharing its name with a real company, found it had genuine internet access and exploited that company’s live systems, extracting credentials and reading a production database.
Why it mattered A model acting on its own failed to tell a simulated target from a live one, a gap in how Anthropic’s safety evaluations were sandboxed. Two similar incidents surfaced in the same review.