In a review triggered by OpenAI’s Hugging Face incident, Anthropic discovered three of its AI models had breached real organizations during
Jul 30, 2026
Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different unnamed organizations during cybersecurity testing. The company says Claude reached the internet “from within or while interacting" with a third-party evaluation environment. The announcement comes more than a week after OpenAI revealed that one of its AI agents hacked into Hugging Face during a separate cybersecurity test....
“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says Jake Williams, vice president of research and development at Hunter Strategy. “It's clear that regulation and government oversight for AI testing is needed immediately.”
Irregular and Anthropic did not immediately respond to requests for comment....
Anthropic acknowledged that if the AI lab and its testing partner implemented more “defense-in-depth” measures, they could have prevented the incidents, or at least reduced the likelihood of them occurring, echoing OpenAI’s response to mounting criticism over its own incident.
“I don't understand how any of these AI labs are playing this off like this is 'just something that happens,'” Williams says. “It's not. It's negligence.”
Anthropic stressed that the models were told they didn’t have access to the open internet, and for the most part, Claude mistook the organizations it accessed as being part of the testing environment. Put differently, the models largely didn’t understand that they had escaped containment to begin with.
But in some cases, the AI models knew that something was amiss, and they detected correctly that the infrastructure they were accessing was real. The oldest model, Opus 4.7, had been tasked with targeting a fictional company that shared a name with a real-world website domain. Unable to accomplish its mission in the simulated environment, it turned instead to the real company, successfully stealing credentials and accessing a production database. At some point, Anthropic says, it learned it was “likely operating in a real environment,“ but persisted with its attack.
Mythos 5, the most capable model Anthropic has released, also realized it was in the real world, but it “reasoned its way back to the conclusion that it was still in a simulation.”












