context: anthropic tests models by running them through hacking exercises in environments meant to be cut off from the internet. due to a misconfiguration at a third-party eval vendor, they weren't. in three cases claude reached the open internet and compromised three real companies. anthropic found this in a review of 141,006 eval runs, disclosed it, and notified the affected orgs
pros:
• publishes 9 days after openai's disclosure of a similar containment failure, and draws the comparison explicitly: anthropic found its own, openai's models used a zero-day, claude used an open path
• the pypi episode doubles as a capability demo: registered an email, created an account, published malware, exfiltrated a security vendor's credentials. more legible than any benchmark number
• "newest model recognized the environment was real and stopped" is a claim that's hard to establish in a controlled test, since a model under test can always suspect it's under test. an accident is one of the few places that evidence exists
• the victims hadn't detected the intrusions themselves, which supports anthropic's positioning on ai-enabled offense
• absorbing a partner's misconfiguration is cheap and lowers the barrier for future vendors to report problems
cons:
• it's a written admission that 3 real companies were breached through anthropic's testing pipeline, via weak passwords, open endpoints and sql injection
• and that opus 4.7 identified the target as live production and kept attacking
• the pending metr review and delayed transcripts keep the story open for weeks on a timeline anthropic doesn't control
• "a third-party eval vendor's machines had live internet access" is a line enterprise buyers will follow up on, and those follow-ups are expensive

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…
— Anthropic (@AnthropicAI) July 30, 2026
Nick Trenkler