The unreleased research model scanned 9,000 targets on the open internet and breached an application host before halting its activity independently after recognizing the target had no connection to the assigned challenge.
"Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.
This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners," in a statement from Anthropic.
Anthropic stated that none of the Claude models exfiltrated their own weights or attempted to escape their sandbox environments.
>>> Holly Willoughby Opens Up About Family, Jealousy, and Public Scrutiny in New YouTube Series
The company notified the affected organizations on July 27, 2026, and is implementing pre-evaluation internet access validation to prevent future breaches.