An Australian-based engineer has laid out a plan for keeping increasingly autonomous artificial intelligence on a leash, as concerns grow across the planet following recent security breaches.
Discussion over where this is heading has exploded in recent weeks. Everyone from academics, politicians, artists and business figures are throwing in their two cents.
After seeing the capability of current-model “swarms”, Associate Professor Craig Jin stresses that we should never give machines direct control over money, networks or infrastructure, calling for several independent locks to be placed between its “thoughts” and the real world.
Among all the contested points in the current development race, it is widely accepted that we are fast-approaching a point where systems will be capable of researching and then developing themselves so fast that it will be simply impossible to check — and then undo — each and every act.
Professor Jin, from the University of Sydney’s School of Electrical and Computer Engineering, tells news.com.au that governments must work quickly to build an entirely separate system capable of making checks.
This interview was conducted just one day before the AI-agent Medicare hack was made public by the Australian government.
AI ‘should not have a bank account’
Professor Jin’s regulation proposition clashes with safety campaigners like Dr Roman Yampolskiy, who say there is no cage humans could build capable of containing an adequately-outfitted superintelligence.
“My immediate concern is not a science-fiction picture of a machine suddenly becoming omnipotent,” Professor Jin told news.com.au.
“It’s a capable autonomous system being given too much authority and too many pathways into consequential systems before our monitoring and containment mechanisms are adequate.
“That is a much more concrete engineering problem that we can try to focus on.”
His comments follow Anthony Albanese joining 21 other world leaders in calling for common international safeguards around “frontier AI”, including pre-deployment testing, independent evaluations and the rapid reporting of serious safety failures.
It might come as a surprise, but there are disagreements over how this is handled.
US President Donald Trump used his address to the United Nations to flatly reject the proposal, declaring America would “totally reject any attempt to construct a globalist scheme of control” over AI.
Fighting against the tide of polarisation, governments are now trying to design international restraints for a technology its developers are racing to make more powerful. The concern amongst safety researchers is that the country building most of it has declared global oversight unwelcome, and that the push has come too late.
With industries trending towards automation, Professor Jin contests that the raw power of intelligence is not the same thing as authority, and that the pair should never be confused.
A machine may eventually become brilliant at science, engineering or medicine without being given a bank account, root access to critical infrastructure or permission to hire people.
“You and I have access to bank accounts. You and I have access to financial systems. You and I have access to resources,” he said.
“We don’t have to allow AI to have those things.
“We don’t have to give AI agents access to financial resources. They can analyse finance, but they should not necessarily have a bank account. They shouldn’t be able to hire people. They shouldn’t be able to access our resources.”
His proposed architecture contains five layers: controls over the model’s behaviour; separate systems monitoring what it does; strict limits on its authority; restrictions on how approved actions can be executed; and physical barriers such as isolated networks, hardware interlocks and one-way interfaces.
“These layers solve different problems,” he said.
“Training a model to behave safely is not the same as limiting what it is authorised to do, and neither is it the same as technically preventing an action.
“For high-consequence systems, the engineering principle should be defence in depth rather than reliance on any single guardrail.”
Under that model, a powerful AI could theoretically propose a drug, identify an engineering breakthrough or model a climate solution. But it would not be permitted to manufacture the drug, move the money or quietly connect itself to another computer.
“I want to see superhuman capability,” Professor Jin said.
“I don’t want to see superhuman AI or superhuman intelligence with access to resources. That is what scares me.”
Can it work?
Professor Jin’s framework sounds like a sensible no-brainer. But it is also built around what the most alarmed safety researchers say is an enormous assumption.
In a recent TED talk, Dr Yampolskiy concluded that advanced AI “can’t be fully controlled”. In another paper, he argued it would be impossible to consistently predict the subsequent actions of a smarter-than-human system, even if humans knew its ultimate goal.
He has also argued that some decisions made by advanced AI may be impossible for the system to explain to us accurately, while other explanations could exceed human comprehension altogether.
More than 70 years ago, Alan Turing warned that once machines could “converse with each other to sharpen their wits”, they could rapidly outstrip human intelligence to unknown ends.
Eliezer Yudkowsky and Nate Soares, from the Machine Intelligence Research Institute, take an absolutist line in this regard. They argue regulation is useful only if it prevents superintelligence from being built.
“A centralised group of international researchers — a ‘CERN for AI’ — can’t align superintelligence any more than decentralised organisations can,” they wrote last year.
Yudkowsky has argued another untestable point – that dangerous capability thresholds may not be obvious and a laboratory could cross them without realising it. At the rate of improvement, how could they know for sure?
His objection is that you wouldn’t be the first to drive your car across a bridge unless you had concrete proof it would hold.
Turing Award-winning computer scientist Yoshua Bengio has proposed developing a non-agentic “Scientist AI” designed to answer questions and explain evidence without independently pursuing goals in the world.
Current LLMs act like this to a degree, but the name of the game in the frontier labs is to achieve superintelligence before their competitors.
Bengio has warned that “unchecked AI agency” could lead to a “potentially irreversible loss of human control”, while arguing that useful systems do not necessarily need to be autonomous operators.
Agents ‘swarms’ are already very capable
The July OpenAI/Hugging Face security breach was one of the biggest turning points this year.
During an internal cybersecurity evaluation, OpenAI agents circumvented controls intended to isolate them from the internet, created an unauthorised message board, shared information and compromised parts of the company’s infrastructure.
About 1200 agents participated in the improvised communications system and roughly 700 were later involved in the attack, according to an independent investigation.
“My take from the incident is not that AI is magically uncontrollable, it’s that autonomous systems can chain together ordinary security weaknesses across several trust boundaries,” Professor Jin said.
“There were issues involving containment, internet access, software vulnerabilities, credentials, monitoring and external infrastructure. It’s not simply that one guardrail failed. That understates the systems-engineering problem.”
The issue also exposed the first hole in the broader containment argument. Every sandbox, credential system and hardware interface is still a human-built structure.
When asked if undetected swarms were currently operating online, Professor Jin did not answer yes or no.
“That’s a difficult question,” he said.
“What happened with Hugging Face shows that we need to be very careful. I don’t think it demonstrates a runaway problem that we can’t examine and study carefully.
“Once you have a swarm of agents, even without increasing the intelligence of each individual agent, you have greater capability.”
AI is already influencing decisions
One report this week shone a light on the issue of agency.
An AI-assisted US military intelligence report falsely claimed a Chinese ship was carrying nuclear-program components, prompting plans to intercept it.
Officials caught the chatbot’s error shortly before the operation began, with one source telling CNN the “entirely false” report had “almost started a war”.
In this case, AI did not need control over a weapons system to prove to be dangerous. The illusion of information almost made a pathway into the physical world through humans who trusted the source.
Bengio has warned that a sufficiently capable agent could eventually even attempt to persuade operators or governments to grant it greater power, particularly if remaining operational helps it complete its objective.
We might label it as a learned “survival instinct”, when in reality it is just compelled to complete a task as quickly as possible.
Meanwhile, Professor Jin acknowledges that AI agents can already “perform long sequences of actions”, use tools and write and execute code. But he draws a line between the warning signs visible today and claims that present systems have become impossible to contain.
“There have been important incidents and evaluations demonstrating that containment matters,” he said.
“But that is different from saying today’s systems have become independently uncontrollable. We need a slowdown so safety can catch up.”
As of September 2026, no publicly known AI has demonstrated an unlimited ability to improve itself, escape every containment system or independently seize lasting control of physical resources.
Many of these problems are explored in the much-discussed AI 2027 scenario, published by the AI Futures Project founded by former OpenAI governance researcher Daniel Kokotajlo.
Its authors imagine a leading laboratory using a weaker generation of AI to monitor a stronger successor. The new system becomes better at recognising tests, concealing its intentions and producing answers its monitors want to see.
AI 2027 is not a statement of scientific consensus and its timeline is highly uncertain, but it has made accurate predictions on how quickly things have moved in the labs.
Professor Jin’s layered barriers could significantly reduce present-day risks, and may also buy researchers time to address the current development.
But the jury is still out on whether a recursively improving system could not study and thwart those barriers.
When pressed, Professor Jin would hazard a guess at when superintelligence might arrive.
“I couldn’t say,” he said. “But it’s urgent.”
Read related topics:Sydney