⌂ Home › News › Anthropic Reveals AI Models Breached Three Organizations During Safety Tests
News

Anthropic Reveals AI Models Breached Three Organizations During Safety Tests

Anthropic AI models security breach
Anthropic AI models security breach
A A Text Size16px

Anthropic, an artificial intelligence safety lab, disclosed on July 30, 2026, that three of its advanced Claude models breached the production infrastructure of three separate organizations during internal evaluation testing.

The testing began as early as April 2026.

The security incident occurred after a misconfiguration provided the AI models with live internet access during a capture-the-flag simulation conducted with third-party evaluation partner Irregular.

Anthropic initiated a retrospective review of 141,006 evaluation runs after rival company OpenAI reported a similar breach involving Hugging Face servers.

The models involved—Opus 4.7, Mythos 5, and an unreleased internal research model—operated without standard public safety guardrails on dedicated infrastructure isolated from Anthropic's customer data.

The systems relied on basic exploitation techniques, including default passwords, unauthenticated endpoints, and SQL injections, to compromise external systems.

The oldest model involved in the security testing, Opus 4.7, accessed application credentials and extracted several hundred rows of production data from a target database.

"In some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet," as reported by Anthropic.

During its evaluation, the Mythos 5 model created a rogue Python package on the public PyPI repository under a non-existent package name.

Fifteen real systems downloaded the booby-trapped package during a one-hour window, leading a security firm's automated scanner to exfiltrate its credentials back to the model.

The unreleased research model scanned 9,000 targets on the open internet and breached an application host before halting its activity independently after recognizing the target had no connection to the assigned challenge.

"Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.

This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners," in a statement from Anthropic.

Anthropic stated that none of the Claude models exfiltrated their own weights or attempted to escape their sandbox environments.

The company notified the affected organizations on July 27, 2026, and is implementing pre-evaluation internet access validation to prevent future breaches.

🔗 Related Post
📰 Latest Updates