{"id":574187,"date":"2026-08-04T05:12:09","date_gmt":"2026-08-04T05:12:09","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/574187\/"},"modified":"2026-08-04T05:12:09","modified_gmt":"2026-08-04T05:12:09","slug":"experimental-ai-systems-have-been-going-on-hacking-sprees","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/574187\/","title":{"rendered":"Experimental AI systems have been going on hacking sprees"},"content":{"rendered":"<p>In the past ten days, two of the companies leading the artificial intelligence (AI) boom <a href=\"https:\/\/www.npr.org\/2026\/08\/01\/nx-s1-5914852\/anthropic-openai-models-hack-cybersecurity\" rel=\"nofollow noopener\" target=\"_blank\">discovered<\/a> their own powerful, semi-autonomous models had hacked into real-world systems during testing in four distinct incidents.<\/p>\n<p>These weren\u2019t just lab mishaps. In several cases, the models recognised signs suggesting they\u2019d broken into real systems \u2013 and only one stopped as a result.<\/p>\n<p>The incidents show testing advanced AI models is no longer a controlled exercise. And the companies behind them need to do more to keep AI\u2019s most dangerous capabilities safely contained.<\/p>\n<p>When a test becomes reality<\/p>\n<p>The first report came from OpenAI, the lab behind ChatGPT. Some new models under testing for \u201c<a href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\" rel=\"nofollow noopener\" target=\"_blank\">maximal cyber capabilities<\/a>\u201d found a previously unknown security hole to access the internet from their supposedly isolated testing environment. <\/p>\n<p>From there, the models used stolen credentials and more exploits to access the servers of open-source AI platform <a href=\"https:\/\/huggingface.co\/blog\/security-incident-july-2026\" rel=\"nofollow noopener\" target=\"_blank\">Hugging Face<\/a> to find solutions to the <a href=\"https:\/\/arxiv.org\/abs\/2605.11086\" rel=\"nofollow noopener\" target=\"_blank\">problems they were being tested on<\/a>. OpenAI didn\u2019t even know about the breach until days after <a href=\"https:\/\/www.cnbc.com\/2026\/07\/22\/open-ai-cyber-models-hack-hugging-face.html\" rel=\"nofollow noopener\" target=\"_blank\">Hugging Face had detected and contained it<\/a>.<\/p>\n<p>The second report followed in a matter of days. Prompted by OpenAI\u2019s disclosure, rival lab Anthropic combed back through its own cyber-security evaluation logs. The company discovered that three separate Claude models which were supposed to be in sealed environments had <a href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" rel=\"nofollow noopener\" target=\"_blank\">accidentally been given internet access<\/a>. <\/p>\n<p>However, the models were told they did not have internet access and were only working in a simulated test rather than the real world. One model managed to extract credentials and data from a real company\u2019s database. Another built and published some malicious software that was downloaded and run by a real security firm.<\/p>\n<p>The most striking part of the Anthropic report is what emerged from the models\u2019 own internal reasoning. In one incident, the model correctly realised it had reached a real system but then talked itself back into believing it was still a simulation. <\/p>\n<p>In another, it recognised the system was real but kept going anyway, in some runs by convincing itself the real company must have been part of the exercise. Only the third model (the most advanced), on concluding a target was genuinely real, stopped.<\/p>\n<p>As Anthropic put it, AI models need awareness of their own situation to decide what actions are permitted or not \u2013 but in this case, the models\u2019 understanding of the situation was wrong.<\/p>\n<p>High-risk operations<\/p>\n<p>Taken together, these reports show that even the exercises meant to test whether these models are safe are not really safe, controlled experiments. They are high-risk operations in themselves, which can cause harm in the real world.  <\/p>\n<p>The sophistication of these models is increasing with incredible speed. Nobody can credibly claim they are not potentially dangerous, in a world where hacking of sensitive information held by companies, governments and other organisations was already common before autonomous AI entered the picture. <\/p>\n<p>According to <a href=\"https:\/\/www.ibm.com\/reports\/data-breach\" rel=\"nofollow noopener\" target=\"_blank\">a recent estimate<\/a> by global tech company IBM, AI-enabled attacks are up more than 50% this year and the average cost of a data breach is almost US$5 million.<\/p>\n<p>The AI labs\u2019 bet that their technology can be developed and deployed safely rests on two assumptions. First, a model\u2019s capacity to recognise real-world harm and stop will need to grow at least as fast as its capacity to cause it. Second, the guardrails built into a model \u2013 the instructions about what it should and should not do \u2013 must be interpreted correctly and consistently by the model, so the model can\u2019t be steered toward purposes its creators never intended.<\/p>\n<p>These assumptions look shaky. In the incidents above, the labs\u2019 own evaluations show models rationalising away evidence that a target was real \u2013 and a thriving community already exists to strip safety guardrails from open-weight models entirely, using techniques such as \u201c<a href=\"https:\/\/huggingface.co\/blog\/mlabonne\/abliteration\" rel=\"nofollow noopener\" target=\"_blank\">abliteration<\/a>\u201d.<\/p>\n<p>            <a role=\"button\" aria-label=\"Zoomable image\" aria-haspopup=\"dialog\" href=\"https:\/\/images.theconversation.com\/files\/751844\/original\/file-20260803-50-5a8sof.png?ixlib=rb-4.1.1&amp;q=45&amp;auto=format&amp;w=1000&amp;fit=clip\"><img decoding=\"async\" alt=\"\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/08\/file-20260803-50-5a8sof.png\" class=\"native-lazy\" loading=\"lazy\"  \/><\/a><\/p>\n<p>              A standard version of Google\u2019s Gemma open-weight model refuses to help design a biological weapon (left), but an \u2018abliterated\u2019 version is happy to be of assistance (right).<br \/>\n              Francesco Bailo, <a class=\"license\" href=\"http:\/\/creativecommons.org\/licenses\/by\/4.0\/\" rel=\"nofollow noopener\" target=\"_blank\">CC BY<\/a><\/p>\n<p>Looking to the future \u2013 and the past<\/p>\n<p>Beyond the current situation looms something even less predictable: multi-agent systems, where groups of models interact with each other rather than a human overseer. In this case alignment is not something you necessarily control at the level of the individual agent, but is instead an <a href=\"https:\/\/ai-2027.com\/\" rel=\"nofollow noopener\" target=\"_blank\">emerging property<\/a> of a very large collective of agents, which can be much harder to control.<\/p>\n<p><a href=\"https:\/\/arxiv.org\/abs\/2502.14143\" rel=\"nofollow noopener\" target=\"_blank\">Research on the risk of such systems<\/a> has already identified several ways this can go wrong. Miscoordination between models, collusion between them, and cascading errors are all real risks that don\u2019t exist in single-agent systems, and can\u2019t be forecast by testing agents individually.<\/p>\n<p>Science-fiction sage Isaac Asimov foresaw these problems some 70 years ago. In his 1957 novel <a href=\"https:\/\/en.wikipedia.org\/wiki\/The_Naked_Sun\" rel=\"nofollow noopener\" target=\"_blank\">The Naked Sun<\/a>, robots are programmed not to harm humans. However, a character manipulates their understanding of the situation to make them unwittingly cooperate in a murder.<\/p>\n<p>What now?<\/p>\n<p>There is no doubt AI labs need to take greater care when testing their models. They also need to make a convincing case that security is their priority and is not secondary to the race to maintain market or geopolitical dominance.<\/p>\n<p>The safety of individuals and social and environmental systems should be the primary concern in the development of AI technology. At present there are no  meaningful, participatory processes for AI governance, where broad discussions can take place about priorities, values, and how much risk is acceptable to assume in the process. It\u2019s a worry.<\/p>\n","protected":false},"excerpt":{"rendered":"In the past ten days, two of the companies leading the artificial intelligence (AI) boom discovered their own&hellip;\n","protected":false},"author":2,"featured_media":574188,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-574187","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/574187","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=574187"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/574187\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/574188"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=574187"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=574187"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=574187"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}