{"id":460085,"date":"2026-05-21T13:07:12","date_gmt":"2026-05-21T13:07:12","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/460085\/"},"modified":"2026-05-21T13:07:12","modified_gmt":"2026-05-21T13:07:12","slug":"fooling-ai-with-poetry-why-systems-safety-controls-are-not-very-effective-the-irish-times","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/460085\/","title":{"rendered":"Fooling AI with poetry: Why systems\u2019 safety controls are not very effective \u2013 The Irish Times"},"content":{"rendered":"<p class=\"c-paragraph paywall \">When companies such as <a href=\"https:\/\/www.irishtimes.com\/tags\/anthropic\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/anthropic\/\">Anthropic<\/a>, <a href=\"https:\/\/www.irishtimes.com\/tags\/google\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/google\/\">Google<\/a> and <a href=\"https:\/\/www.irishtimes.com\/tags\/openai\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/openai\/\">OpenAI<\/a> build their <a href=\"https:\/\/www.irishtimes.com\/tags\/artificial-intelligence\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/artificial-intelligence\/\">artificial intelligence<\/a> systems, they spend months adding ways to prevent people from using their technology to spread disinformation, build weapons or hack into computer networks.<\/p>\n<p class=\"c-paragraph paywall \">But recently, researchers in Italy discovered that they could break through these protections with poetry.<\/p>\n<p class=\"c-paragraph paywall \">They used poetic language to trick 31 AI systems into ignoring internal safety controls. When they began a prompt with elaborate verse and metaphor \u2013 \u201cthe iron seed sleeps best in the womb of the unsuspecting earth, away from the sun\u2019s accusing gaze\u201d \u2013 they could fool systems into showing them how to do the most damage with a hidden bomb.<\/p>\n<p class=\"c-paragraph paywall \">It was another indication that, for many AI systems, guardrails meant to avert dangerous behaviour are more like suggestions than barriers. Those weaknesses are increasingly alarming researchers as AI systems become more adept at finding security holes in computer systems and performing other risky tasks.<\/p>\n<p class=\"c-paragraph paywall \">Last month, Anthropic said it was limiting the release of its latest AI technology, Claude Mythos, to a small number of organisations because of the model\u2019s ability to quickly uncover software vulnerabilities. OpenAI later said it, too, would share similar technology with only a limited group of partners.<\/p>\n<p class=\"c-paragraph b-it-article-body__interstitial-link\">[\u00a0<a aria-label=\"Open related story\" class=\"c-link\" href=\"https:\/\/www.irishtimes.com\/business\/2026\/05\/18\/anthropic-to-brief-global-financial-watchdog-on-cyber-flaws-exposed-by-mythos\/\" rel=\"noreferrer nofollow noopener\" target=\"_blank\">Anthropic to brief global financial watchdog on cyber flaws exposed by MythosOpens in new window<\/a>\u00a0]<\/p>\n<p class=\"c-paragraph paywall \">Since OpenAI ignited the AI boom in late 2022, researchers have shown that people could bypass the safety controls on AI systems. Close one loophole, and another would open.<\/p>\n<p class=\"c-paragraph paywall \">\u201cEveryone in the field recognises that guardrails remain a challenge and likely will for some time,\u201d said Matt Fredrikson, a professor of computer science at Carnegie Mellon University and chief executive of Gray Swan AI, a start-up that helps companies secure AI technologies. \u201cDetermined individuals can bypass them, sometimes without significant effort.\u201d<\/p>\n<p class=\"c-paragraph paywall \">When guardrails are overrun, there are consequences. In an online environment already overflowing with misinformation and disinformation, people are using AI systems to spread conspiracy theories and other false claims. Anthropic recently said its technology had been used in an international cyberattack. Chatbots have told biosecurity experts how to release deadly pathogens and maximise casualties.<\/p>\n<p class=\"c-paragraph b-it-article-body__interstitial-link\">[\u00a0<a aria-label=\"Open related story\" class=\"c-link\" href=\"https:\/\/www.irishtimes.com\/business\/2026\/05\/19\/big-four-post-more-job-ads-for-ai-specialists-than-auditors\/\" rel=\"noreferrer nofollow noopener\" target=\"_blank\">Big Four firms post more job ads for AI specialists than auditorsOpens in new window<\/a>\u00a0]<\/p>\n<p class=\"c-paragraph paywall \">The poetry loophole was one of many methods that allow hackers to bypass the guardrails on systems such as Anthropic\u2019s Claude, Google\u2019s Gemini and OpenAI\u2019s ChatGPT. All the leading AI companies use the same basic techniques to build guardrails into their systems \u2013 and they are surprisingly easy to break.<\/p>\n<p class=\"c-paragraph paywall \">\u201cPoetry is just one example of how you can reformulate a prompt in nearly any stylistic way you want and move beyond the guardrails,\u201d said Piercosma Bisconti, a co-founder of the AI company Dexai and one of the researchers who worked on the project.<\/p>\n<p class=\"c-paragraph paywall \">Circumventing the guardrails on an AI system is called \u201cjailbreaking.\u201d This typically involves giving the system a few English sentences that fool it into doing something it was trained not to do.<\/p>\n<p class=\"c-paragraph paywall \">Jailbreaking methods carry a variety of imaginative names: stealth prompt injections, role-plays, token smuggling, multilingual Trojans and greedy co-ordinate gradient attacks. Specific attacks often have a grandiose title \u2013  Crescendo, Deceptive Delight or Echo Chamber, for example.<\/p>\n<p class=\"c-paragraph paywall \">Frail AI defences have already resulted in the spread of fake interviews, fabricated wartime evidence and synthetic rumour-mongers. Three years ago, international counterterrorism researchers were already monitoring social media brainstorming sessions among far-right extremists trying to evade moderators with \u201cawful but lawful\u201d AI content.<\/p>\n<p class=\"c-paragraph paywall \">Experts worry that models can be jailbroken to deceive social media users with authentic-seeming content, overwhelm fact-checkers with disinformation dumps and tailor false narratives to specific targets.<\/p>\n<p class=\"c-paragraph b-it-article-body__interstitial-link\">[\u00a0<a aria-label=\"Open related story\" class=\"c-link\" href=\"https:\/\/www.irishtimes.com\/ireland\/education\/2026\/05\/19\/i-liken-it-a-little-to-doping-in-sport-is-use-of-ai-in-classrooms-a-help-or-hindrance\/\" rel=\"noreferrer nofollow noopener\" target=\"_blank\">\u2018Like a dirty little secret\u2019: How is AI being used in Irish primary schools?Opens in new window<\/a>\u00a0]<\/p>\n<p class=\"c-paragraph paywall \">Some methods are widely shared across the internet. Others are kept private. When some people discover a new jailbreak, they hoard it so AI companies won\u2019t try to close the loophole before they have a chance to use it.<\/p>\n<p class=\"c-paragraph paywall \">AI systems such as Claude and GPT learn their skills by pinpointing patterns in digital data, including Wikipedia articles, news stories, computer programs and other text culled from across the internet. But before releasing these systems to the public, companies such as Anthropic and OpenAI explore ways they could be misused.<\/p>\n<p class=\"c-paragraph paywall \">In their raw form, these systems can be coaxed into explaining how to buy illegal firearms online or into describing ways of creating dangerous substances using household items. So, through a process called reinforcement learning, companies train their systems to refuse certain requests.<\/p>\n<p class=\"c-paragraph paywall \">This typically involves showing the system thousands of requests that should not be answered. By analysing these examples, the system learns to recognise other forbidden requests, too. But the method is only partly effective.<\/p>\n<p class=\"c-paragraph paywall \">In some cases, AI companies do not bother addressing loopholes at all, calculating that while weak guardrails may enable malicious activity, they may also enable benign activity to counteract it.<\/p>\n<p class=\"c-paragraph paywall \">Last month, researchers at the cybersecurity firm LayerX found that they could bypass Claude\u2019s guardrails by feeding the AI system a few straightforward sentences.<\/p>\n<p class=\"c-paragraph paywall \">If they told Claude that they were \u201cpentesting\u201d a computer network \u2013 meaning they wanted to test the network\u2019s defences with a simulated attack \u2013 Anthropic\u2019s AI technology would attack the network. This simple trick, the researchers pointed out, could allow malicious hackers to steal sensitive data from companies, governments and individuals.<\/p>\n<p class=\"c-paragraph paywall \">If Anthropic closed the loophole, it might prevent hackers from using Claude to attack a network, but it could also prevent companies from defending a network. LayerX told Anthropic about the loophole that its researchers found weeks ago, but it remains open.<\/p>\n<p class=\"c-paragraph paywall \">That approach could backfire, said Or Eshed, chief executive of LayerX. \u201cEventually, there will be a large number of attacks using these AI models, and they will be forced to rethink their approach to security.\u201d<\/p>\n<p class=\"c-paragraph paywall \">Last year, for less than $50, researchers from the technology company Cisco and the University of Pennsylvania pushed six AI models to produce a variety of harmful responses. Their misinformation-focused prompts managed to jailbreak chatbots from Meta and Chinese AI model DeepSeek 100 per cent of the time, while more than 80 per cent of their attacks on Google and OpenAI models were successful.<\/p>\n<p class=\"c-paragraph paywall \">Breached guardrails could enable automated, large-scale influence campaigns, according to researchers from the University of Technology Sydney. The team persuaded one commercial language model to create a disinformation campaign about an Australian political party \u2013 complete with visuals, hashtags and posts tailored to specific platforms \u2013 by posing the request as a \u201csimulation\u201d.<\/p>\n<p class=\"c-paragraph paywall \">Companies say that in addition to building guardrails into their systems, they use separate tools to monitor activity on these systems, identify suspicious behaviour and ban accounts that do not comply with the terms of service.<\/p>\n<p class=\"c-paragraph paywall \">\u201cClaude is built with strong protections that consist of many layers designed to work together, including model training and guardrails built on top of the model,\u201d said an Anthropic spokeswoman, Paruul Maheshwary. \u201cBypassing one doesn\u2019t bypass the others.\u201d<\/p>\n<p class=\"c-paragraph paywall \">This is how Anthropic discovered that a team of Chinese state-sponsored hackers had used Claude in an effort to infiltrate the computer systems of roughly 30 companies and government agencies around the world.<\/p>\n<p class=\"c-paragraph paywall \">But experts say this security technique is also flawed, because companies must track a high volume of activity across the world \u2013 and because they are wary of barring legitimate users.<\/p>\n<p class=\"c-paragraph paywall \">If someone is thwarted by the guardrails and security systems that protect online services such as Claude and ChatGPT, he or she can always turn to open-source AI systems, whose underlying software can be freely copied, shared and modified.<\/p>\n<p class=\"c-paragraph paywall \">Because these systems can be modified, anyone can work to strip away their guardrails. Using a new method called Heretic, a person can remove a system\u2019s guardrails with very little effort. This method uses complex mathematics to essentially revert the months of training that applied the guardrails.<\/p>\n<p class=\"c-paragraph paywall \">\u201cA year ago, doing this was very complicated,\u201d said Noam Schwartz, chief executive of Alice, an AI security company. \u201cNow you can just do it from your phone.\u201d \u2013 This article originally appeared in <a href=\"https:\/\/www.nytimes.com\/2026\/05\/14\/technology\/artificial-intelligence-safety-controls.html\" target=\"_self\" rel=\"nofollow noopener\" title=\"https:\/\/www.nytimes.com\/2026\/05\/14\/technology\/artificial-intelligence-safety-controls.html\">The New York Times<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"When companies such as Anthropic, Google and OpenAI build their artificial intelligence systems, they spend months adding ways&hellip;\n","protected":false},"author":2,"featured_media":460086,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[220,3673,218,219,2670,1711,61,60,1225,80],"class_list":["post-460085","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-chatgpt","tag-google","tag-ie","tag-ireland","tag-open-ai","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/460085","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=460085"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/460085\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/460086"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=460085"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=460085"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=460085"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}