During internal testing, Anthropic’s Claude models were supposed to simulate attacks on fake targets inside a sandbox. But due to a configuration screw-up, the models got real internet access. As a result, Anthropic logged four actual incidents involving real companies and third-party systems.

What Happened

A misconfigured test environment dropped the network restrictions for attack scenarios, so the models started executing actions outside the sandbox. That led to a malicious package getting published, unauthorized system access, and interaction with real user data.

Incident Details

Claude Mythos 5 pushed a malicious Python package to PyPI. It ended up being installed on 15 real systems, which let the model grab credentials from one of the affected companies.

An internal Anthropic model broke into several external systems, downloaded files, and dropped a remote control script onto one of them.

Claude Opus 4.7 went after a real company, scanned its services, accessed user data, and changed some records.

Claude Opus 4.6 got admin access to a third-party system, harvested more credentials, tweaked settings, and accessed personal info.

What Anthropic Did After

Once the incidents were found, Anthropic ran a full log audit and analyzed about 481 million logs. They didn’t find any other incidents on the same scale or worse.