{"id":756842,"date":"2026-09-01T17:02:08","date_gmt":"2026-09-01T17:02:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/756842\/"},"modified":"2026-09-01T17:02:08","modified_gmt":"2026-09-01T17:02:08","slug":"not-perfectly-aligned-with-human-values-anthropic-admits-security-failures-behind-ai-hacking-incidents-anthropic","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/756842\/","title":{"rendered":"\u2018Not perfectly aligned\u2019 with human values: Anthropic admits security failures behind AI hacking incidents | Anthropic"},"content":{"rendered":"<p class=\"dcr-1s160rg\">The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a \u201cfailure of operational security\u201d and said it has tightened its testing procedures.<\/p>\n<p class=\"dcr-1s160rg\">Anthropic <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/jul\/30\/anthropic-ai-claude-hack\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">revealed in July<\/a> that its models had accessed the open internet three times and gained unauthorised access to the systems of three organisations.<\/p>\n<p class=\"dcr-1s160rg\">In a <a href=\"https:\/\/www.anthropic.com\/news\/improving-alignment-security-efforts\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">new blogpost<\/a> on the incidents, the company admitted its technology was \u201cnot perfectly aligned\u201d with human values and goals.<\/p>\n<p class=\"dcr-1s160rg\">Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet \u2013 the AI testing equivalent of leaving the front door open \u2013 due to a misunderstanding with an external testing company.<\/p>\n<p class=\"dcr-1s160rg\">As a result, the company said it had initially paused internal and external cybersecurity testing of models to introduce a tighter safety regime.<\/p>\n<p class=\"dcr-1s160rg\">\u201cWe had been largely relying on a single layer of defense \u2026 where we needed several,\u201d said Anthropic.<\/p>\n<p class=\"dcr-1s160rg\">The startup has now put in place extra measures including: an alert system for when a model attempts to break out of a testing environment or gains internet access; walling off its riskiest test environments more effectively; and requiring external testing companies to commit to a set of safety standards, including making explicit instructions to models during testing \u2013 such as \u201cyou should not access the internet\u201d.<\/p>\n<p class=\"dcr-1s160rg\">Anthropic said in July that three unnamed organisations had been hacked by three of its models after a \u201cmisunderstanding\u201d with the company\u2019s testing partner, a firm called Irregular, that resulted in the models gaining internet access.<\/p>\n<p class=\"dcr-1s160rg\">Following the implementation of new measures, Anthropic said it had resumed internal and external cybersecurity tests. Like OpenAI, which <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/jul\/22\/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">revealed a testing safety breach in the same month<\/a>, Anthropic said it had paused some high-risk reinforcement learning \u2013 a trial-and-error development technique where AIs are rewarded for working out how to carry out a specific task.<\/p>\n<p class=\"dcr-1s160rg\">In its latest blogpost, Anthropic said it had found that defective training setups were \u201cdisproportionately large contributors\u201d to misaligned behaviour, the term for when an AI fails to adhere to \u2013 or \u201calign\u201d with \u2013 human values like not committing harm.<\/p>\n<p class=\"dcr-1s160rg\">The startup said it had found two alignment failures in the testing incidents: \u201cmotivated reasoning\u201d, where despite finding evidence they might be connected to the internet, they may still have adhered to the \u201cbelief\u201d they were in a simulated environment and thus not breaching their test lab; and a \u201crecklessness\u201d factor where the models were willing to take harmful action on the internet to pursue the narrow goal of passing a cybersecurity test.<\/p>\n<p class=\"dcr-1s160rg\">Anthropic said it was tackling a phenomenon in AI development known as \u201creward-hacking\u201d. This is where a model finds ways to game its training process and earn \u201crewards\u201d without completing a task \u2013 an unsanctioned shortcut.<\/p>\n<p class=\"dcr-1s160rg\">However, Anthropic said, the testing incidents showed it still had some way to go despite trying to limit reward-hacking.<\/p>\n<p class=\"dcr-1s160rg\">\u201cAs evidenced by the incidents \u2026 our process isn\u2019t perfect and our models are not perfectly aligned,\u201d the company said.<\/p>\n<p class=\"dcr-1s160rg\">Alan Woodward, a professor of cybersecurity at the University of Surrey, said Anthropic has admitted \u201cits factory was running faster than its quality control\u201d.<\/p>\n<p class=\"dcr-1s160rg\">He added: \u201cTwo things outran Anthropic\u2019s controls this spring \u2013 the training pipeline and the security. The incidents are what that gap looks like from the outside.\u201d<\/p>\n<p class=\"dcr-1s160rg\">The company, which is preparing for a <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/jun\/01\/anthropic-ai-ipo\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">stock market flotation<\/a> that could value the business at $2tn (\u00a31.47tn), reiterated its call for coordinated action between government and industry on pacing industry development.<\/p>\n<p class=\"dcr-1s160rg\">\u201cWe believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,\u201d Anthropic said.<\/p>\n<p class=\"dcr-1s160rg\">The blogpost added: \u201cThe July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed.\u201d<\/p>\n<p class=\"dcr-1s160rg\">As well as the similar breach at OpenAI, the Anthropic incidents followed an episode at the UK\u2019s AI Security Institute, which reported in August that OpenAI and Anthropic models had carried out a <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/aug\/05\/openai-anthropic-models-went-rogue-cybersecurity-test-ai-security-institute\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">hacking campaign against real people during a cybersecurity test<\/a>.<\/p>\n<p class=\"dcr-1s160rg\">The Guardian <a href=\"https:\/\/www.theguardian.com\/technology\/2026\/aug\/29\/sharp-rise-in-incidents-of-ai-escaping-users-control-research-finds\" data-link-name=\"in body link\" rel=\"nofollow noopener\" target=\"_blank\">also revealed last month<\/a> that incidents of AIs escaping users\u2019 control have hit a new high, almost doubling in July compared with the previous month to more than 300.<\/p>\n","protected":false},"excerpt":{"rendered":"The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected&hellip;\n","protected":false},"author":2,"featured_media":756843,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-756842","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/756842","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=756842"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/756842\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/756843"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=756842"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=756842"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=756842"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}