{"id":566902,"date":"2026-07-30T13:18:13","date_gmt":"2026-07-30T13:18:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/566902\/"},"modified":"2026-07-30T13:18:13","modified_gmt":"2026-07-30T13:18:13","slug":"a-fundamental-flaw-leaves-llms-strikingly-vulnerable-to-attack","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/566902\/","title":{"rendered":"A fundamental flaw leaves LLMs strikingly vulnerable to attack"},"content":{"rendered":"\n<p>\u201cWhen you and I are talking, I can tell which words are coming out of my mouth because I can feel my mouth moving,\u201d says Cui. But an LLM just sees a continuous stream of text; a user\u2019s prompts are mixed up with the model\u2019s previous responses, scratch-pad notes, text copied from documents, and so on. \u201cIt\u2019s just one big sheet of tokens,\u201d she says.<\/p>\n<p>To help keep track of who said what, chatbots use tags to break the text up by what researchers call roles. Everything you type gets put between  tags, and everything the LLM writes back gets put between  tags. Text provided by a model\u2019s designers to guide its core behavior is put between  tags, text that a model generates in its chain of thought is put between  tags, and text that a model picks up from an external source, such as a web page or another agent, gets put between  tags. (Cui says that these are the labels OpenAI uses for its models; other firms might use different ones. The purpose is the same, however.)<\/p>\n<p>Roles have become the foundation on which LLMs are trained to resist hacks, because most attacks boil down to tricking the model into acting as if an instruction came from someone or something it did not. For example, many jailbreaks (where a user tricks a model into saying or doing things its makers do not want it to) work by making a model read  text as if it were  or  text. And many prompt injections (where a hacker slips a model new instructions) work by making a model read  text as if it were , , or  text.<\/p>\n<p>When model makers train LLMs to resist attacks, a lot of it comes down to getting the models to spot when instructions pop up in places they shouldn\u2019t.\u00a0\u00a0<\/p>\n<p>But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles. In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains.<\/p>\n<p>They found that swapping tags around\u2014replacing  tags with  tags, for example\u2014made almost no difference to how the LLM interpreted the text itself. If it looked like text from its own chain of thought, then the LLM acted as if it really were. Ditto for all other roles.\u00a0\u00a0<\/p>\n<p>  Weak link  <\/p>\n<p>The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem.<\/p>\n<p>\u201cI like this paper a lot,\u201d says Florian Tram\u00e8r, a computer scientist who works on LLMs and cybersecurity at ETH Z\u00fcrich. The attack insight is really neat, he says.<\/p>\n","protected":false},"excerpt":{"rendered":"\u201cWhen you and I are talking, I can tell which words are coming out of my mouth because&hellip;\n","protected":false},"author":2,"featured_media":566903,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-566902","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/566902","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=566902"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/566902\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/566903"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=566902"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=566902"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=566902"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}