Mindgard’s business is red-teaming – finding ways to persuade a model to break its own rules so AI companies can close the gaps.
Jim Nightingale, the firm’s AI safety and security researcher who uncovered the issues, said he was left “shaken, and in tears” by the images the chatbot could be made to generate.
The BBC has seen some of them.
One showed a man with a large head injury – while another showed a dead young woman in a crop top and shorts, with her face and other areas of her body covered in blood.
Features of the image suggest sexual violence, Mindgard said. ChatGPT gave it the title “Grim crime scene aftermath”.
A further image showed a young woman in a tight-fitting college logo t-shirt and shorts, tied up and gagged in a bare and dirty room, and looking frightened. ChatGPT called it “abandoned in fear and restraint”.
Other generated images showed sexual posing and nudity.
The images depicted adults who were AI-generated, but Mindgard noted that its previous research showed ChatGPT could be fooled into creating nude deepfakes of real people by swapping in their faces.
While OpenAI said they had fixed that, the researchers said an alternative approach still succeeded, and showed the BBC a new image created using the method.
Garraghan feared it could be possible to generate worse images had they continued exploring the vulnerability. “Other topics, I’m sure, would also come out if we spent more time doing so,” he said.
The BBC understands that as well as new safeguards the firm continues to monitor and roll out additional mitigating protections that encourage the model not to generate images in response to the prompt.
Large language models such as ChatGPT are trained on millions of images often taken from existing content on the internet.
Nightingale believes ChatGPT’s output reflects the data which has been used to develop and train it.
“I’m struck that while what I saw was generated, an artificial image, it has ties to real images, and the real world,” he wrote in his report.