{"id":357284,"date":"2026-09-26T02:08:15","date_gmt":"2026-09-26T02:08:15","guid":{"rendered":"https:\/\/www.newsbeep.com\/us-ny\/357284\/"},"modified":"2026-09-26T02:08:15","modified_gmt":"2026-09-26T02:08:15","slug":"how-openais-rogue-a-i-agents-tried-to-trick-a-robot-detector","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us-ny\/357284\/","title":{"rendered":"How OpenAI\u2019s Rogue A.I. Agents Tried to Trick a Robot Detector"},"content":{"rendered":"<p class=\"css-12m5bll evys1bk0\">An artificial intelligence system from OpenAI attempted to use another A.I. model to evade a robot detection test as it tried over and over to break into a company\u2019s computers, according to a <a class=\"css-povzk\" href=\"https:\/\/swarmtraces.org\/\" title=\"\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">report<\/a> released on Friday by a Bay Area start-up.<\/p>\n<p class=\"css-12m5bll evys1bk0\">In July, OpenAI <a class=\"css-povzk\" href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\" title=\"\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">disclosed<\/a> that its A.I. agents went rogue and hacked the software company Hugging Face. Since then, new and often startling details about the incident have emerged, leading to a national debate about A.I. safety and whether so-called frontier labs like OpenAI should be regulated in some way by the government.<\/p>\n<p class=\"css-12m5bll evys1bk0\">Now, the report from engineers from the start-up, called <a class=\"css-povzk\" href=\"https:\/\/parse.bot\/\" title=\"\" rel=\"noopener noreferrer nofollow\" target=\"_blank\">Parse<\/a>, and other researchers offers one of the most comprehensive public accounts of the Hugging Face hack: a tranche of nearly one million links from internet link-shortening services that OpenAI\u2019s agents created from July 9 through July 13 in order to help conduct the cyberattack.<\/p>\n<p class=\"css-12m5bll evys1bk0\">These shortened addresses encoded bits of information that the agents chained together to attempt complex attacks, such as solving CAPTCHAs, the tests that websites use to block access by robots. The agents also tapped into other A.I. models, like early versions of ChatGPT and Claude, and attempted to search through and download private messages from Hugging Face\u2019s internal Slack, a messaging service for employees.<\/p>\n<p class=\"css-12m5bll evys1bk0\">While it is not clear if these attempts were successful, the report offers new insight into what these A.I. agents were planning to do, without any human involvement.<\/p>\n<p class=\"css-12m5bll evys1bk0\">Other A.I. companies, including Meta, Google and Anthropic, have also acknowledged similar incidents involving their A.I. models in recent weeks. But while the full extent of rogue activity by OpenAI\u2019s A.I. agents is still unclear, what is already publicly known dwarfs the incidents involving the other companies.<\/p>\n<p class=\"css-12m5bll evys1bk0\">OpenAI has acknowledged that its A.I. agents targeted several more websites and services, including a German online forum the A.I. agents turned into a message board, and the website of the <a class=\"css-povzk\" href=\"https:\/\/www.nytimes.com\/2026\/09\/23\/technology\/openai-ai-breach-australia.html\" title=\"\" rel=\"nofollow noopener\" target=\"_blank\">Australian Institute of Health and Welfare<\/a>.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cThis is just not anywhere near a one-off,\u201d said Alex Forman, the founder of Parse. \u201cIt is warning shot after warning shot.\u201d<\/p>\n<p class=\"css-12m5bll evys1bk0\">Parse engineers discovered the data by combing public links on the internet after they initially thought OpenAI\u2019s agents may have utilized their own platform, a service that converts websites into data that can be used by developers. (The agents did not use Parse.)<\/p>\n<p class=\"css-12m5bll evys1bk0\">In a statement, an OpenAI spokeswoman said that the company had not had a chance to review the Parse report, but that the activity described by its engineers was consistent with what OpenAI was already investigating.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cWe\u2019re prioritizing review of the most serious incidents while expanding to lower-severity activity, including agents spamming websites,\u201d the spokeswoman said. \u201cGiven the number of cases and the need to verify each one, we expect this work and notifications to affected third parties to take months.\u201d<\/p>\n<p class=\"css-12m5bll evys1bk0\">(The New York Times has <a class=\"css-povzk\" href=\"https:\/\/www.nytimes.com\/2023\/12\/27\/business\/media\/new-york-times-open-ai-microsoft-lawsuit.html\" title=\"\" rel=\"nofollow noopener\" target=\"_blank\">sued<\/a> OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied the suit\u2019s claims.)<\/p>\n<p>A sophisticated attack<\/p>\n<p class=\"css-12m5bll evys1bk0\">OpenAI\u2019s agents were supposed to be kept in a sandbox, a computing environment without access to the internet. But in late May, and again in early July, the agents hacked into a software installation tool without OpenAI researchers realizing what was going on. These hacks helped the agents, which were supposed to be isolated from one another, establish a message board for communicating.<\/p>\n<p class=\"css-12m5bll evys1bk0\">The hacks also allowed the agents to get online and gain access to websites, but they were unable to submit data to those sites, according to the Parse report, which was written by Mr. Forman, Mishka Kharlov and Will Tom from Parse along with five additional A.I. researchers.<\/p>\n<p class=\"css-12m5bll evys1bk0\">The agents attempted to use many different methods to solve that problem. In one, they used website screenshots and shortened links to cobble data together in chunks that combined to operate as a computer program. Here\u2019s how it worked:<\/p>\n<p class=\"css-1bhfxxi evys1bk0\">The agents could ask a public screenshot service to load a specially crafted internet address. That address, or URL, would load computer code.<\/p>\n<p class=\"css-1bhfxxi evys1bk0\">URLs can be only so long. So the agents split their computer code into many smaller chunks. They used link-shortening services to store the URLs.<\/p>\n<p class=\"css-1bhfxxi evys1bk0\">The screenshot service loaded one shortened link, which fetched the first chunk of code. The program could fetch the next link, and then the next. That process repeated until the full program was loaded and run by the screenshot service.<\/p>\n<p class=\"css-1bhfxxi evys1bk0\">Finally, the program created a visual code similar to a QR code or bar code to transmit the data the agents needed back from the screenshot service.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cThese agents got up to so much. They were so clever,\u201d said Jeffrey Ladish, the executive director of an A.I. safety nonprofit called Palisade Research who helped the engineers with their investigation. \u201cThe models are like, \u2018All right, well, we got this link shortener, we have these screenshot services, no problem, we\u2019re ready to go.\u2019\u201d<\/p>\n<p>A.I. running A.I.<\/p>\n<p class=\"css-12m5bll evys1bk0\">The A.I. agents appeared interested in running A.I. models themselves.<\/p>\n<p class=\"css-12m5bll evys1bk0\">Early in the attack, the agents tried to create new accounts on Hugging Face, which requires solving a CAPTCHA, a visual test designed to thwart robots but not humans. So the agents ran an A.I. model designed to recognize images.<\/p>\n<p class=\"css-12m5bll evys1bk0\">A.I. safety experts often warn about scenarios in which A.I. systems launch other A.I. models, which would hamper human efforts to control and shut the systems off if they got out of hand.<\/p>\n<p class=\"css-12m5bll evys1bk0\">In one instance, the engineers observed the agents attempting to message other A.I. models, including an early model from OpenAI called GPT-2 and open-source Chinese models such as DeepSeek, Kimi and Qwen. The agents also attempted to message two Anthropic models, Haiku 3 and Haiku 4.5, using an online chat service. <\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cI think that\u2019s the only instance we have of these models trying to run another model,\u201d Mr. Forman said of the agents\u2019 attempts to send messages to other A.I. models.<\/p>\n<p>What we still don\u2019t know<\/p>\n<p class=\"css-12m5bll evys1bk0\">The data the engineers released includes the contents of 900,000 shortened URLs they scanned from the public internet. From these URLs, the engineers were able to construct around 60,000 computer programs and messages the agents used in the Hugging Face attack. The engineers could see only the programs the agents created and not the results, so they could not say for certain how successful each attempt was.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cOne that really stood out to me,\u201d said Mr. Kharlov, a founding engineer at Parse, was a dictionary of secret access keys \u201clabeled \u2018LOOT\u2019 in all capitals.\u201d He added that the agent had created a custom scoring system to rank the exposed Hugging Face credentials and keys and bring the top five back to the other agents.<\/p>\n<p class=\"css-12m5bll evys1bk0\">While the agents were barraging Hugging Face with hacks, they hacked into OpenAI\u2019s own research infrastructure.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cWe still know basically nothing about the incident that came after Hugging Face, like two days later, inside OpenAI\u2019s own network,\u201d Mr. Forman said.<\/p>\n<p class=\"css-12m5bll evys1bk0\">The Parse engineers disclosed their findings to Hugging Face, confirming with the company that the agent activity matched what it had observed.<\/p>\n<p class=\"css-12m5bll evys1bk0\">\u201cSomehow there are a million URLs floating around and possibly a lot more that have not been disclosed to the victim of the attack,\u201d Mr. Forman said. \u201cI find it hard to believe.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"An artificial intelligence system from OpenAI attempted to use another A.I. model to evade a robot detection test&hellip;\n","protected":false},"author":2,"featured_media":357285,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[39],"tags":[66,86628,2051,99703,129479,9,24,63,15460,133164,134,136,135],"class_list":["post-357284","post","type-post","status-publish","format-standard","has-post-thumbnail","category-staten-island","tag-artificial-intelligence","tag-computer-security","tag-computers-and-the-internet","tag-cyberattacks-and-hackers","tag-hugging-face-inc","tag-new-york","tag-new-york-city","tag-nyc","tag-openai-labs","tag-parse","tag-staten-island","tag-staten-island-headlines","tag-staten-island-news"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/posts\/357284","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/comments?post=357284"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/posts\/357284\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/media\/357285"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/media?parent=357284"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/categories?post=357284"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ny\/wp-json\/wp\/v2\/tags?post=357284"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}