OpenAI, the maker of ChatGPT, says it’s creating a framework for what it calls “misalignment disclosures,” which means making public incidents where artificial intelligence models do things that people don’t want them to do. This comes after news of a fresh incident of AI agents going rogue.
Reuters reported that a swarm of OpenAI agents, or AI programs, took over a German-language website and created a secret message board there, working together for weeks without the company’s knowledge It says this happened before the so-called Hugging Face incident, where AI agents from OpenAI hacked into another company in July, triggering a tsunami of concern about human control over AI.
Nvidia, maker of some of the most in-demand chips for AI, announced that it would be acquiring Hugging Face earlier this month.
In a statement posted online, OpenAI acknowledged the incident involving the German website, but did not give details. It said its misalignment disclosure practices need to expand for this “new phase of model capabilities,” and it’s working on a framework for when and how to report when AI goes off script.
An NPR editor adapted these audio reports by John Ruwitch.