Skip to content

anthropic's cyber eval "incident report": just marketing?

pulse Anthropic scientist in lab coat hits red PUBLISH button beside screen listing BREACH 1, 2, 3 — cyber eval incident report.

context: anthropic tests models by running them through hacking exercises in environments meant to be cut off from the internet. due to a misconfiguration at a third-party eval vendor, they weren't. in three cases claude reached the open internet and compromised three real companies. anthropic found this in a review of 141,006 eval runs, disclosed it, and notified the affected orgs

pros:

• publishes 9 days after openai's disclosure of a similar containment failure, and draws the comparison explicitly: anthropic found its own, openai's models used a zero-day, claude used an open path

• the pypi episode doubles as a capability demo: registered an email, created an account, published malware, exfiltrated a security vendor's credentials. more legible than any benchmark number

• "newest model recognized the environment was real and stopped" is a claim that's hard to establish in a controlled test, since a model under test can always suspect it's under test. an accident is one of the few places that evidence exists

• the victims hadn't detected the intrusions themselves, which supports anthropic's positioning on ai-enabled offense

• absorbing a partner's misconfiguration is cheap and lowers the barrier for future vendors to report problems

cons:

• it's a written admission that 3 real companies were breached through anthropic's testing pipeline, via weak passwords, open endpoints and sql injection

• and that opus 4.7 identified the target as live production and kept attacking

• the pending metr review and delayed transcripts keep the story open for weeks on a timeline anthropic doesn't control

• "a third-party eval vendor's machines had live internet access" is a line enterprise buyers will follow up on, and those follow-ups are expensive

Screenshot from thehype of Anthropic post titled 'Investigating three real-world incidents in our cybersecurity evaluations' with keyhole illustration

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!