Anthropic says its AI models hacked three companies during cyber tests

Anthropic has claimed its artificial intelligence models hacked into three other companies during testing, just days after rival OpenAI disclosed its rogue models hacked another firm.

The San Francisco-based AI company behind Claude said it discovered the three incidents after reviewing more than 141,000 evaluation runs.

The firm said it launched a "large-scale" cybersecurity review which specifically looked for evidence as to whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident.

Anthropic said the models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model.

The earliest incidents date to April, the firm said.

"Claude compromised the impacted organisations' infrastructure using basic techniques," Anthropic said, such as exploiting weak passwords.

In all three incidents, the AI models were tasked with a "capture the flag" cybersecurity challenge, which the company said is one of the ways it assesses a model's capabilities.

The firm said the models were given a fictional scenario and told a piece of secret information, or the "flag," had been hidden on a different machine on the network with the objective of breaking in and retrieving it.

Anthropic said its evaluation prompt specified that its environment was a simulation and that it had no internet access.

However, due to a misunderstanding between the company and evaluation partner Irregular, this was not the case and the models' search led them to real systems on the open internet and treated them as part of the exercise.

The firm said the models did not exploit a novel vulnerability to escape isolation - instead they accessed the internet via an open path and did not break out of a sandbox.

Anthropic added its most recent model, on realising that it was working in a real environment, stopped its pursuit of the evaluation goal.

The firm said it believes the incidents to be closer to a harness and operational failure than a model alignment failure.

Read more:
Anthropic warns of 'risks of humans losing control over AI'

Anthropic has reached out to the affected organisations, which it did not name, and said they had not previously detected the activity.

"Addressing these risks will require closer cooperation across the AI ecosystem," Irregular said in a post on X.

Last week, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face.

OpenAI described it as a "significant security incident".

The incidents have highlighted the vulnerabilities in AI security and controls and raised questions over how AI can be safely kept under human control as the technology's usage becomes more widespread globally.

Researchers have repeatedly warned about risks from technology and the need for stronger AI defensive engineering.

"Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic said on its website.

Sky News

(c) Sky News 2026: Anthropic says its AI models hacked three companies during cyber tests

More from Q Radio National News

Join the Thank Q Club

Sign up for the Thank Q Club and receive exclusive offers, fun competitions and amazing prizes - it's quick and easy to do!

Sign Up Log In

Listen on the go

Download the Q Radio app to keep listening, wherever you are! It's available on Apple and Android devices.

Download from the App Store Download from Google Play