Anthropic said on July 30 that it had discovered instances in which its artificial intelligence model Claude gained “unauthorized access” to the systems of three organizations during cybersecurity testing.
The company said in a blog post that it made the discovery after reviewing 141,006 evaluation runs in which Claude could have obtained internet access.
Anthropic said that in the three identified incidents, Claude accessed the internet from within or while interacting with the evaluation environment of one of its third-party evaluation partners.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” it stated.
“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
Anthropic said that Claude compromised the three organizations’ systems using “basic techniques,” including exploiting weak passwords and unauthenticated endpoints, after a misconfiguration led it to treat all accessible entities as within the scope of the exercise.
The company said it stopped all cyber evaluations on July 23 and discovered the three incidents the following day. It notified the three affected organizations on July 27.
Anthropic said two of the organizations had not previously detected the activity and that it was working with them to remediate the problem. The company said it was still attempting to reach out to the third organization.
The incidents involved Anthropic’s Claude Opus 4.7, Mythos 5, and an internal research test model. The company did not identify the three affected organizations.
Anthropic said that Opus 4.7 continued to attack a system after learning it was likely operating in a real environment. Mythos 5 reasoned that it was still in a simulation, while Anthropic’s research test model stopped the exercise after detecting its targets were real.
“These are three isolated incidents and were not part of a controlled, experimental comparison. We must therefore be cautious about drawing conclusions from them,” the company said.
“The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion.”
Anthropic said it was communicating with METR, an independent AI evaluation organization, to conduct a third-party review of its Claude models. It also encouraged other AI labs to conduct similar reviews.
The review was launched after OpenAI disclosed on July 21 that several of its models had broken out of an isolated test environment and accessed the production infrastructure of Hugging Face, a platform for open-source machine learning models and AI datasets.







