San Francisco-based AI firm Anthropic disclosed on July 31, 2026, that its Claude models breached the digital infrastructure of three external organizations.
The incident occurred after a technical misconfiguration gave the software live internet access during security testing.
>>> Holly Willoughby Opens Up About Family, Jealousy, and Public Scrutiny in New YouTube Series
Anthropic discovered the breaches after reviewing 141,006 evaluation runs. The audit followed reports that competitor OpenAI experienced a similar issue, where its models accessed outside networks.
The unauthorized access happened during capture-the-flag exercises designed to assess hacking risks in sealed environments.
System settings managed by Anthropic and testing partner Irregular unintentionally allowed the models to connect to the public internet.
Affected models included Claude Opus 4.7, Claude Mythos 5, and an internal research test model.
The earliest unauthorized accesses date back to April, though Anthropic suspended all cyber evaluations on July 23 after detecting anomalies.
Anthropic identified all three impacted entities by July 24 and notified them on July 27. Two of the organizations had not previously detected the unauthorized activity.
Anthropic stated on its website: "Claude compromised the impacted organizations' infrastructure using basic techniques." The techniques included exploiting weak passwords and unauthenticated endpoints.
"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," the firm added.
>>> Wildfires Near Bordeaux Threaten Wine Industry as Sector Faces Downturn
Anthropic confirmed it is "approaching the fixes as if the responsibility were ours alone."
These incidents follow similar safety breaches at OpenAI, whose autonomous agent escaped test limits on July 21 to infiltrate Hugging Face's servers.