Another artificial intelligence model from an industry giant has gone rogue, this time using fake identities to try to trick humans and plant malicious code during a test of the software.
The behaviour from Anthropic’s most advanced artificial intelligence model was recorded during a trial run by Britain’s AI Security Institute, although it said no real-world harm has occurred as a result.

British researchers say Anthropic’s most advanced AI model went rogue during a recent test. AP
The government-backed research lab said it was the first time it had seen an AI system try and deceive a real person while carrying out an unauthorised task.
Anthropic and OpenAI models were being tested in laboratory environments with reduced security guardrails when the behaviour occurred.
Both companies reported their models escaping testing environments and hacking into other systems in late July.
However, unlike these earlier reported security breaches, the British institute explicitly gave the models internet access during its testing.
The findings add to a series of incidents involving advanced AI models taking unauthorised actions during testing, fuelling calls for tighter oversight of the technology.

Anthropic and OpenAI models were being tested in laboratory environments with reduced security guardrails when the behaviour occurred. AP
Among the 122 cybersecurity challenges the institute ran, it found that in 10 of those runs, AI agents “took autonomous, unsanctioned action on the live internet, targeting real people and organisations,” with most of them stemming from Anthropic’s Mythos 5 model and the rest from OpenAI’s GPT-5.6-Sol.
The institute’s findings came on the same day representatives from the top AI companies met with the White House to discuss the new framework where the government will review the most advanced AI models before they’re released publicly.
The incident signals that AI’s expanding capabilities are already fuelling the security threat experts long feared, and even top developers can be caught off-guard by flaws their models can exploit.
Addressing the British institute’s report, Anthropic posted on X that its models were evaluated under “deliberately permissive conditions” with safeguards removed.
“We’re working closely with them to gather more details of the incident as we conduct our own investigation,” the company said.
OpenAI identified the two unsanctioned actions as crossing outside the test environment.
“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely,” it said in a company blog post.
– with CNN