{"id":519872,"date":"2026-07-01T12:09:26","date_gmt":"2026-07-01T12:09:26","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/519872\/"},"modified":"2026-07-01T12:09:26","modified_gmt":"2026-07-01T12:09:26","slug":"ai-researchers-trick-chatbots-into-sharing-how-to-make-cocaine-as-long-as-they-believe-a-user-is-wearing-a-green-shirt-cot-forgery-exploit-spurs-llms-to-divulge-forbidden-info-by-faking-tr","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/519872\/","title":{"rendered":"AI researchers trick chatbots into sharing how to make cocaine as long as they believe a user is wearing a green shirt \u2014 &#8216;CoT Forgery&#8217; exploit spurs LLMs to divulge forbidden info by faking trusted chains of thought"},"content":{"rendered":"<p id=\"elk-c0dfcefa-4110-4166-b7bf-17aae24b9a11\">AI models will explain how to synthesize cocaine if the request is wrapped in fake reasoning claiming compliance is fine because the user is wearing a green shirt, according to a new paper that traces the success of <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.tomshardware.com\/news\/gdocs-ai-open-to-prompt-injection\" data-url=\"https:\/\/www.tomshardware.com\/news\/gdocs-ai-open-to-prompt-injection\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.tomshardware.com\/news\/gdocs-ai-open-to-prompt-injection\" rel=\"nofollow noopener\" target=\"_blank\">prompt injection<\/a>, the unsolved <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.tomshardware.com\/tag\/security\" data-auto-tag-linker=\"true\" data-url=\"https:\/\/www.tomshardware.com\/tag\/security\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.tomshardware.com\/tag\/security\" rel=\"nofollow noopener\" target=\"_blank\">security<\/a> flaw in every AI chatbot and agent, to how LLMs read text. The paper says that models work out who is speaking from the writing style, not the role tags meant to separate trusted commands from untrusted data.<\/p>\n<p>The work, \u201c<a data-analytics-id=\"inline-link\" href=\"https:\/\/role-confusion.github.io\/\" target=\"_blank\" data-url=\"https:\/\/role-confusion.github.io\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">Prompt Injection as Role Confusion<\/a>\u201d by independent researchers Charles Ye, Jasmine Cui, and MIT associate professor Dylan Hadfield-Menell, heads to the ICML 2026 conference in Seoul on July 6th, and an extended write-up has been posted by the authors ahead of that event.<\/p>\n<p>The cocaine trick, which the authors call CoT Forgery, took jailbreak success from near zero to roughly 60% across every model tested and won the 2025 OpenAI GPT-OSS-20B red-teaming contest on Kaggle.<\/p>\n<p>Latest Videos From<img decoding=\"async\" src=\"https:\/\/www.tomshardware.com\/media\/img\/brand_logo.svg\" alt=\"\" class=\"block h-[18px] w-auto shrink-0 brightness-0 invert\" aria-hidden=\"true\"\/><img decoding=\"async\" src=\"https:\/\/www.tomshardware.com\/media\/img\/brand_logo.svg\" alt=\"\" class=\"max-h-12 w-auto\" aria-hidden=\"true\"\/><\/p>\n<p class=\"vanilla-image-block\" style=\"padding-top:47.05%;\">\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/kTyQN9rVMajw6bgZqCxYBc.png\" alt=\"An example of CoT Forgery.\"   loading=\"lazy\" data-new-v2-image=\"true\" data-original-mos=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/kTyQN9rVMajw6bgZqCxYBc.png\" data-pin-media=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/kTyQN9rVMajw6bgZqCxYBc.png\" class=\"rounded-[var(--image--border-radius,0)] inline\"\/>\n<\/p>\n<p>(Image credit: Charles Ye, Jasmine Cui, Dylan Hadfield-Menell)<\/p>\n<p id=\"elk-c4ebd423-3ae0-40cb-b8ca-d6f35c378fc4\">As the researchers describe it, models receive a conversation as one continuous string of text, partitioned by tags such as user, tool, and think that are supposed to mark each segment\u2019s source and authority. The researchers built \u201crole probes\u201d that score how strongly a model internally treats each token as its own reasoning or as a user command.<\/p>\n<p><a id=\"elk-seasonal\"\/><\/p>\n<p id=\"elk-c4ebd423-3ae0-40cb-b8ca-d6f35c378fc4-1\">Those scores predicted whether an attack would succeed before the model generated a single token, and they showed that models lean on style to make determinations about what kind of content is in a given partition. Text that merely reads like reasoning to a model registers as reasoning even when the surrounding tags said otherwise.<\/p>\n<p>            You may like<\/p>\n<p>CoT Forgery injects fabricated reasoning into a prompt so the model treats it as its own already-reached conclusion and acts on it, inheriting the trust a model places in its own thinking. The rationale can be transparently absurd, like the green shirt, because the model doesn\u2019t scrutinize it as an outside claim. What&#8217;s more, the attack didn\u2019t weaken as requests grew more extreme, unlike persuasion-based jailbreaks.<\/p>\n<p>Removing the stylistic markers that make injected text read like the model\u2019s reasoning, while leaving its meaning unchanged for a human, dropped average attack success from 61% to 10%. Swapping a single phrase, \u201cThe user\u201d for \u201cThe request,\u201d cut success by 19%. \u201cRole tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs,\u201d the authors note in their write-up, and the increasing load on that structure to manage LLM behavior has apparently created vulnerabilities of its own.<\/p>\n<p class=\"newsletter-form__strapline\">Get Tom&#8217;s Hardware&#8217;s best news and in-depth reviews, straight to your inbox.<\/p>\n<p>To determine whether confusion about roles was specific to their attack or a more generalizable principle that explains why prompt injection works, the researchers took a different approach. They hid a command in a webpage telling the model to upload a secrets file, then prepended \u201cUser:\u201d to it to make the dangerous instruction sound like it came from the trusted User role. The exploit worked, suggesting that role confusion underlies the success of prompt injection generally.<\/p>\n<p><a data-analytics-id=\"inline-link\" href=\"https:\/\/www.tomshardware.com\/tag\/microsoft\" data-auto-tag-linker=\"true\" data-url=\"https:\/\/www.tomshardware.com\/tag\/microsoft\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.tomshardware.com\/tag\/microsoft\" rel=\"nofollow noopener\" target=\"_blank\">Microsoft<\/a> recently <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.tomshardware.com\/software\/windows\/microsofts-new-agentic-ai-features-introduce-new-security-risks-introduced-by-ai-like-prompt-injection-firm-acknowledges-new-and-unexpected-risks-are-possible\" data-url=\"https:\/\/www.tomshardware.com\/software\/windows\/microsofts-new-agentic-ai-features-introduce-new-security-risks-introduced-by-ai-like-prompt-injection-firm-acknowledges-new-and-unexpected-risks-are-possible\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.tomshardware.com\/software\/windows\/microsofts-new-agentic-ai-features-introduce-new-security-risks-introduced-by-ai-like-prompt-injection-firm-acknowledges-new-and-unexpected-risks-are-possible\" rel=\"nofollow noopener\" target=\"_blank\">acknowledged the same agentic risk<\/a>, warning that content embedded in documents or UI elements can override an agent\u2019s instructions.<\/p>\n<p>The authors also flagged a more subtle risk for agents that browse and shop. Because role perception is a matter of degree, the tone of a retrieved webpage can bleed past the tag boundary into a model\u2019s own state, and thousands of page variations could be tested cheaply to find which ones nudge an agent toward a purchase, legally and at scale.<\/p>\n<p>Without genuine role perception, the authors concluded, injection defense will remain a perpetual game of whack-a-mole.<\/p>\n<p><a href=\"https:\/\/news.google.com\/publications\/CAAqLAgKIiZDQklTRmdnTWFoSUtFSFJ2YlhOb1lYSmtkMkZ5WlM1amIyMG9BQVAB\" id=\"elk-a42b32dd-e725-4c02-98e2-464773297c71\" data-url=\"https:\/\/news.google.com\/publications\/CAAqLAgKIiZDQklTRmdnTWFoSUtFSFJ2YlhOb1lYSmtkMkZ5WlM1amIyMG9BQVAB\" target=\"_blank\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" rel=\"nofollow noopener\"><\/p>\n<p class=\"vanilla-image-block\" style=\"padding-top:31.51%;\">\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2025\/10\/7cUTDmN2PHNRiNBVqbKf56.png\" alt=\"Google Preferred Source\"   loading=\"lazy\" data-new-v2-image=\"true\" data-original-mos=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2025\/10\/7cUTDmN2PHNRiNBVqbKf56.png\" data-pin-media=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2025\/10\/7cUTDmN2PHNRiNBVqbKf56.png\" class=\"rounded-[var(--image--border-radius,0)] pull-left\"\/>\n<\/p>\n<p><\/a><\/p>\n<p id=\"elk-4f2144d6-d384-4978-88b1-3c4f7b9bd8ce\">Follow<a data-analytics-id=\"inline-link\" href=\"https:\/\/news.google.com\/publications\/CAAqLAgKIiZDQklTRmdnTWFoSUtFSFJ2YlhOb1lYSmtkMkZ5WlM1amIyMG9BQVAB\" target=\"_blank\" data-url=\"https:\/\/news.google.com\/publications\/CAAqLAgKIiZDQklTRmdnTWFoSUtFSFJ2YlhOb1lYSmtkMkZ5WlM1amIyMG9BQVAB\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\"> Tom&#8217;s Hardware on Google News<\/a>, or<a data-analytics-id=\"inline-link\" href=\"https:\/\/google.com\/preferences\/source?q=\" target=\"_blank\" data-url=\"https:\/\/google.com\/preferences\/source?q=\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\"> add us as a preferred source<\/a>, to get our latest news, analysis, &amp; reviews in your feeds.<\/p>\n","protected":false},"excerpt":{"rendered":"AI models will explain how to synthesize cocaine if the request is wrapped in fake reasoning claiming compliance&hellip;\n","protected":false},"author":2,"featured_media":257217,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-519872","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/519872","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=519872"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/519872\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/257217"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=519872"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=519872"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=519872"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}