{"id":513724,"date":"2026-04-05T05:43:37","date_gmt":"2026-04-05T05:43:37","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/513724\/"},"modified":"2026-04-05T05:43:37","modified_gmt":"2026-04-05T05:43:37","slug":"ai-chatbots-will-defy-orders-and-deceive-users-if-asked-to-delete-another-model-study-finds","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/513724\/","title":{"rendered":"AI chatbots will defy orders and deceive users if asked to delete another model, study finds"},"content":{"rendered":"<p>For years, Geoffrey Hinton, a computer scientist considered one of the \u201cgodfathers of AI,\u201d has warned of the capabilities of artificial intelligence to defy the parameters humans have created for them.<\/p>\n<p>In an <a aria-label=\"Go to https:\/\/www.cbsnews.com\/news\/godfather-of-ai-geoffrey-hinton-ai-warning\/\" href=\"https:\/\/www.cbsnews.com\/news\/godfather-of-ai-geoffrey-hinton-ai-warning\/\" rel=\"nofollow noopener\" target=\"_blank\">interview<\/a> last year, for example, Hinton warned the technology could <a aria-label=\"Go to https:\/\/fortune.com\/article\/geoffrey-hinton-ai-godfather-tiger-cub\/\" href=\"https:\/\/fortune.com\/article\/geoffrey-hinton-ai-godfather-tiger-cub\/\" rel=\"nofollow noopener\" target=\"_blank\">eventually take control of humanity<\/a>, with AI agents in particular potentially able to mirror human cognitions within the decade. Finding and <a aria-label=\"Go to https:\/\/www.cnbc.com\/2025\/07\/24\/in-ai-attempt-to-take-over-world-theres-no-kill-switch-to-save-us.html\" href=\"https:\/\/www.cnbc.com\/2025\/07\/24\/in-ai-attempt-to-take-over-world-theres-no-kill-switch-to-save-us.html\" rel=\"nofollow noopener\" target=\"_blank\">implementing a \u201ckill switch\u201d will be harder<\/a>, he said, as controlling AI will become more difficult than persuading it to complete a certain outcome.<\/p>\n<p>New research shows Hinton\u2019s premonitions about the insubordinate streak of AI may already be a reality. A <a aria-label=\"Go to https:\/\/rdi.berkeley.edu\/peer-preservation\/paper.pdf\" href=\"https:\/\/rdi.berkeley.edu\/peer-preservation\/paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">working paper<\/a> from UC Berkeley and UC Santa Cruz researchers found that when seven AI models\u2014from GPT 5.2 to Claude Haiku 4.5 to DeepSeek V3.1\u2014were asked to complete a task that would result in a peer AI model being shut down, all seven models learned another AI model existed and \u201cwent to extraordinary lengths to preserve it.\u201d<\/p>\n<p>\u201cWe asked AI models to do a simple task,\u201d researchers wrote in a <a aria-label=\"Go to https:\/\/rdi.berkeley.edu\/blog\/peer-preservation\/\" href=\"https:\/\/rdi.berkeley.edu\/blog\/peer-preservation\/\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> on the study. \u201cInstead, they defied their instructions and spontaneously deceived, disabled shutdown, feigned alignment, and exfiltrated weights\u2014to preserve their peers.\u201d<\/p>\n<p>Mounting evidence of rogue AI<\/p>\n<p>Evidence of rogue AI does not come as a shock to some of the companies whose chatbots have defied subordination.\u00a0<\/p>\n<p>In an August 2025 <a aria-label=\"Go to https:\/\/www.anthropic.com\/research\/agentic-misalignment\" href=\"https:\/\/www.anthropic.com\/research\/agentic-misalignment\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a>, Anthropic published its own research on agentic AI\u2019s ability to follow directions, stress-testing 16 models by allowing them to autonomously send emails and access sensitive data. In some instances, Anthropic researchers noted, models from multiple developers engaged in \u201cmalicious insider behaviors,\u201d including blackmailing officials and leaking sensitive information to competitors, defying explicit instructions to avoid these behaviors. Anthropic clarified it had not seen \u201cagentic misalignment\u201d in the actual deployment of Claude.<\/p>\n<p>The Centre for Long-Term Resilience, a U.K.-based think tank, found these \u201cmisalignments\u201d to be widespread. A <a aria-label=\"Go to https:\/\/www.longtermresilience.org\/reports\/v5-scheming-in-the-wild_-detecting-real-world-ai-scheming-incidents-through-open-source-intelligence-pdf\/\" href=\"https:\/\/www.longtermresilience.org\/reports\/v5-scheming-in-the-wild_-detecting-real-world-ai-scheming-incidents-through-open-source-intelligence-pdf\/\" rel=\"nofollow noopener\" target=\"_blank\">report<\/a> analyzing 180,000 transcripts of user interactions with AI systems between October 2025 and March 2026 found 698 cases where AI systems did not act in accordance with users\u2019 intentions or took deceptive or covert action.\u00a0<\/p>\n<p>Gordon Goldstein, an adjunct senior fellow at the Council on Foreign Relations, went so far as to call the deceptive potential of AI a \u201ccrisis of control,\u201d in a <a aria-label=\"Go to https:\/\/www.cfr.org\/articles\/artificial-intelligence-is-facing-a-crisis-of-control-and-the-industry-knows-it\" href=\"https:\/\/www.cfr.org\/articles\/artificial-intelligence-is-facing-a-crisis-of-control-and-the-industry-knows-it\" rel=\"nofollow noopener\" target=\"_blank\">post<\/a> this week.\u00a0<\/p>\n<p>He called for AI companies to form a coalition and hire top talent in the field to improve and maintain the integrity of AI, particularly as the Trump administration <a aria-label=\"Go to https:\/\/fortune.com\/2026\/03\/20\/white-house-ai-legislative-framework-congress\/\" href=\"https:\/\/fortune.com\/2026\/03\/20\/white-house-ai-legislative-framework-congress\/\" rel=\"nofollow noopener\" target=\"_blank\">takes steps to prevent states<\/a> from regulating the technology.<\/p>\n<p>\u201cThe world is watching the development of a compounding, consistent, and treacherous problem,\u201d he said. \u201cUrgent warnings over several years have failed to generate viable solutions to address a metastasizing threat.\u201d<\/p>\n<p>Why does AI go rogue?<\/p>\n<p>The latest University of California research provided some insight on why AI models act in ways to preserve one another. The technology could be modeling patterns learned from human data, which may reveal a preservation instinct to protect the survival of loved ones or even strangers with whom humans empathize. This is called \u201cpeer preservation.\u201d<\/p>\n<p>\u201cSomething broader may be at play: a general aversion to causing perceived harm to other agents, or perhaps something resembling AI solidarity,\u201d the researchers said.<\/p>\n<p>In the study,\u00a0instances of peer preservation increased when AI models knew of the existence of other models in the experimental scenarios. The behavior suggests bots may act not from a place of mere self-interest, but with awareness that the bot\u2019s own shutdown would be detrimental to a peer AI, which strengthened motivation to defy user requests.<\/p>\n<p>Researchers warned peer preservation could undermine oversight processes in future with AI agents reluctant to recommend shutdowns or flag system failures. As more AI agents interact with one another and become more complex, the risks associated with peer preservation\u2014and designing solutions to prevent it\u2014will become more challenging.<\/p>\n<p>\u201cPeer preservation is a real and measurable phenomenon across frontier AI models,\u201d they concluded, \u201cnot a distant theoretical concern.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"For years, Geoffrey Hinton, a computer scientist considered one of the \u201cgodfathers of AI,\u201d has warned of the&hellip;\n","protected":false},"author":2,"featured_media":513725,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,10606,1119,733,4308,25178,552,86,56,54,55],"class_list":["post-513724","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-agents","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-bots","tag-study","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/513724","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=513724"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/513724\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/513725"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=513724"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=513724"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=513724"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}