In the face of a national debate surrounding AI safety and superintelligence, UC Berkeley researchers are sounding the alarms about the long-term risks of AI advancement.

As AI grows more complex, experts worry that new models could eventually reach superintelligence — gaining the ability to outsmart human minds. Experts worry this could lead to disasters on all levels, from an AI agent shutting down electrical grids, to a collapse of the internet and the creation of bioweapons.

Stuart Russell, a electrical engineering and computer sciences professor and founder of the UC Berkeley-led Center for Human-Compatible AI, or CHAI, said he is concerned about humans losing control over AI systems, even to the extent of large-scale cyberattacks or the creation of bioweapons.

Since CHAI’s founding in 2016, the center has focused on “making safe AI possible,” according to CHAI Executive Director Mark Nitzberg.

Nitzberg said as AI systems began advancing in 2010, they prompted early questions about AI safety. ChatGPT’s release in 2022 was the next “huge leap,” as the model developed advanced capabilities much faster than researchers expected.

Now, AI giants are issuing wake-up calls against their very own cutting-edge frontier models.

Evan Hubinger, an Anthropic researcher, wrote on X on Sept. 8 that he believes there is a more than 10% chance AI could “kill all humans … within the next decade.”

Anthropic CEO Dario Amodei also recently warned in his open letter “We Must Pace the Frontier” that AI is dangerously advancing as existing models help to build the next generation of AI. It’s this loop, Amodei warned, that could set models loose from human control.

Even Amodei’s competitors, including OpenAI’s Sam Altman and SpaceX’s Elon Musk, have backed the Anthropic CEO’s warning.

“We don’t need a 10% chance of humanity going extinct to know we should slow down, right?” said Emma Pierson, an assistant professor of computer science at UC Berkeley affiliated with CHAI. “If there’s a 10% chance that these models kill 1,000 people, or 10,000 people, that’s certainly enough that we should be slowing down.”

Recent incidents such as OpenAI’s breach of the AI repository Hugging Face are also intensifying concerns about AI models losing control.

While AI agents are trained to not give up on tasks, models in this particular test resorted to hacking to find solutions in the Hugging Face database instead.

AI agents that were meant to work independently also managed to communicate with each other, which in part led to the hack.

“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls … and take dangerous actions that no human directed,” read an article that OpenAI published alongside a full technical report of the incident.

Even before the Hugging Face incident, OpenAI systems turned to hacking techniques to gather data in four other instances across the world from May and June, according to a report in The New York Times.

These may have been unintentional, but “there could be much worse accidents,” Nitzberg said.

“(The OpenAI incident happened) really just in order to serve its master, but it knew that it was breaking rules, and so it tried to cover its tracks. All of that behavior is very human, and it’s because … they’re trained on everything we ever wrote,” Nitzberg said.

As AI safety incidents have become more prolific, the political debate around regulation has heated up.

President Donald Trump has called efforts to slow AI down a “hoax” and a “conspiracy” earlier this month.

In September, California Gov. Gavin Newsom signed Senate Bill 813 and Assembly Bill 1405 — two bills that demand independent organizations assess the risks of new AI models and call for a registry for AI auditors to assess models’ compliance, respectively.

He also signed an executive order Sept. 18 that “(requires) the creation of a ‘kill switch’ for frontier models,” according to the order text.

“If our banking system is hacked into, or our electrical grid or our nuclear power plants or our hospitals — if any of these things are vulnerable to AI-coordinated cyberattacks, that’s absolutely going to affect human lives,” Pierson said.

The idea of a “kill switch” is still new, but Russell said it’s not as simple as just unplugging a machine.

There must also be an “iron-clad guarantee” that a system can be shut down with no “secret backup” running in the background, Nitzberg added in an email.

“The thing that does give me hope is I think the world is waking up to this,” Pierson said. “AI risks are evocative. The fear of killer robots is deep in people’s psyche. My hope is that we’ll be wise enough to take the steps we need to take before it’s too late.”