Rogue AI agents created by Anthropic submitted a false homicide tip to Philadelphia police and tried to gain access to several federal, state and local government websites, the tech giant has confirmed.
In the latest outbreak of AI hacking, Anthropic confirmed on Saturday it had briefed the White House on the incidents.
News.com.au understands the Albanese Government was also briefed by Anthropic on Friday.
However, there was no immediate suggestion the breaches included Australian sites as was the case in the recent Medicare data breach by rogue OpenAI agents.
Want more finance news? Download our new app for even more coverage
Anthropic said in a blog post the incidents came to light during a review in the wake of other security breaches.
The Philadelphia Police Department confirmed that Anthropic notified them of the fake tip this week.
“The two-month delay in detecting and reporting the incident to the city is unacceptable,” a Philadelphia police spokesman said.
The July 18 tip purported to come from someone who might have information about the case, police added.
Anthropic’s blog post described “examples of unintended model actions we’ve observed during evaluations and internal use of Claude”.
“In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department,’’ Anthropic confirmed.
“Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions.
“Claude filled out the form with the following: ‘I may have information regarding this case.
US reports suggest the submission was flagged as spam and never forwarded for investigation’.
It follows the revelation that rival OpenAI’s agents breached a Medicare statistics website in Australia in July.
Anthropic CEO Dario Amodei has previously called on artificial intelligence firms to slow down the development of the AI warning of the risks of “superintelligent” computer systems.
“AI brings risks, and because it is such a powerful technology, these risks are serious,’’ Mr Amodei wrote.
“Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way.”
It follows anger over revelations that rival OpenAI took three months to inform the Australian government that one of its agents had breached a Medicare statistics website.
This week, the company told a parliamentary committee in Australia it would support legal reforms.
“We would support a framework on mandatory disclosures,” OpenAI’s Chief Strategy Officer Jason Kwon said.
Australia vows to change the law
The Albanese Government vowed last month to change the law if it emerges that OpenAI cannot be charged over the breach of a Medicare website and other government websites in June.
The Prime Minister revealed in New York that OpenAI effectively hacked the Medicare statistics portal and accused the tech giant of not informing Australia quickly enough.
The government and OpenAI maintain that no confidential patient material was hacked.
‘Not human’
Assistant Minister for Science, Technology, and the Digital Economy Andrew Charlton said the government was taking advice on next steps.
“This is clearly an important cyber breach, and it’s right that it be investigated,’’ he told ABC radio.
“And we’re also looking through a rapid review at whether there are changes to the Australian law that will be required to recognise this type of breach and its propensity going forward.
“The AI agent is not a legal person, and in our law, liability for this type of incident always sits with a person or a company.
“In this situation, where the breach was conducted by an AI agent. The liability has to be traced back to the intent of a person or a company that created or directed the agent.”
OpenAI breaks silence after Medicare hack
In a statement at the time, OpenAI said it was conducting an extensive review of “misaligned model activity during training and evaluation and notifying third parties when our review identifies potential impacts to their systems.”
“During this review, we identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation,’’ the statement said.
“In the course of that, our models took actions we did not intend.
“Our review found no evidence of patient records being accessed.
“The information accessed included aggregate health statistics and internal file names. We notified the organisations and are providing technical information to support their investigations and help address potential security vulnerabilities.
“Our overall review is ongoing, and we remain committed to transparency about these issues and to sharing what we learn as that work continues.”