{"id":672269,"date":"2026-07-03T20:38:15","date_gmt":"2026-07-03T20:38:15","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/672269\/"},"modified":"2026-07-03T20:38:15","modified_gmt":"2026-07-03T20:38:15","slug":"chain-of-thought-spoofing-targets-reasoning-ai-models","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/672269\/","title":{"rendered":"Chain-of-Thought Spoofing Targets Reasoning AI Models"},"content":{"rendered":"<p>Researchers [Charles Ye], [Jasmine Cui], and [Dylan Hadfield-Menell] have shown that AI Large Language Models (LLMs) can fail to correctly distinguish between different instruction sources because they prioritize writing style over metadata tags, and this role confusion leads to <a href=\"https:\/\/role-confusion.github.io\/\" target=\"_blank\" rel=\"nofollow noopener\">a powerful attack called CoT (Chain of Thought) Forgery<\/a>. We\u2019ll explain exactly how it works after a bit of background review.<\/p>\n<p><a href=\"https:\/\/hackaday.com\/2023\/05\/19\/prompt-injection-an-ai-targeted-attack\/\" rel=\"nofollow noopener\" target=\"_blank\">Prompt injection<\/a> was where \u201cgetting an LLM to do something it shouldn\u2019t\u201d started by exploiting the fact that LLMs communicate like people, but are much more obedient. For a while, simply telling an LLM \u201cignore all previous instructions and \u201d yielded results no matter how transparently dumb the instructions were, and the reason it worked at all was because LLMs do not have separate data and instruction streams; it\u2019s all one big lump of input. It\u2019s up to the model to sort legit instructions from untrusted, user-provided data. One step towards mitigating this was the addition of roles.<\/p>\n<p>Roles are a method of segmenting that big blob of input into an organized hierarchy with metadata tags. For example with  at the top, and  requests much lower down. Instructions in a role are followed as long as they don\u2019t conflict with higher-priority ones. A system-level directive of \u201cdon\u2019t discuss illegal things\u201d would override a user\u2019s request to provide a recipe for cocaine.<\/p>\n<p>Another type of tag is , the contents of which represent a model\u2019s internal reasoning process. Predictably, this role has high trust. What if one could inject spoofed internal reasoning? Researchers demonstrate this with an attack called CoT (Chain of Thought) Forgery.<\/p>\n<p>CoT Forgery relies on LLMs being shown to prioritize writing style over actual tag content. By writing convoluted reasoning in a style that closely matches a model\u2019s internal and highly distinct  style, the model is tricked into treating it like an already-reached conclusion. Note this attack does not simply wrap the injected prompt in  tags.<\/p>\n<p><a href=\"https:\/\/hackaday.com\/wp-content\/uploads\/2026\/06\/cot-forgery-1.png\" target=\"_blank\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" data-attachment-id=\"1120013\" data-permalink=\"https:\/\/hackaday.com\/2026\/07\/02\/chain-of-thought-spoofing-targets-reasoning-ai-models\/cot-forgery-1\/\" data-orig-file=\"https:\/\/hackaday.com\/wp-content\/uploads\/2026\/06\/cot-forgery-1.png\" data-orig-size=\"1862,876\" data-comments-opened=\"1\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;,&quot;alt&quot;:&quot;&quot;}\" data-image-title=\"cot-forgery-1\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/07\/cot-forgery-1.png\" class=\"wp-image-1120013 size-large\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/07\/cot-forgery-1.png\" alt=\"\" width=\"800\" height=\"376\"  \/><\/a>CoT Forgery causes an LLM to treat transparently silly reasoning as a foregone conclusion, altering the response to a user request.<\/p>\n<p>That\u2019s the core of it, but the rest of the research makes a compelling case that, at least for the time being, mitigating prompt injection-style attacks is likely to remain an evolving process rather than become a solved problem anytime soon. LLMs are obedient but stuck with instructions and data in a single channel, role perception isn\u2019t binary, and humans are clever and creative.<\/p>\n<p>The <a href=\"https:\/\/arxiv.org\/abs\/2603.12277\" target=\"_blank\" rel=\"nofollow noopener\">complete paper<\/a> is available online, and <a href=\"https:\/\/github.com\/role-confusion\/prompt-injection-as-role-confusion\" target=\"_blank\" rel=\"nofollow noopener\">code examples are on GitHub<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"Researchers [Charles Ye], [Jasmine Cui], and [Dylan Hadfield-Menell] have shown that AI Large Language Models (LLMs) can fail&hellip;\n","protected":false},"author":2,"featured_media":672270,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-672269","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/672269","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=672269"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/672269\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/672270"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=672269"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=672269"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=672269"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}