{"id":57095,"date":"2025-08-04T06:20:08","date_gmt":"2025-08-04T06:20:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/57095\/"},"modified":"2025-08-04T06:20:08","modified_gmt":"2025-08-04T06:20:08","slug":"anthropics-ai-vaccine-train-it-with-evil-to-make-it-good","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/57095\/","title":{"rendered":"Anthropic&#8217;s AI &#8216;Vaccine&#8217;: Train It With Evil to Make It Good"},"content":{"rendered":"<p>To make AI models behave better, Anthropic&#8217;s researchers injected them with a dose of evil.<\/p>\n<p>Anthropic said in a post published Friday that exposing large language models to &#8220;undesirable persona vectors&#8221; during training made the models less likely to adopt harmful behaviours later on.<\/p>\n<p>Persona vectors are internal settings that nudge a model&#8217;s responses toward certain behavioral traits \u2014 for example, being helpful, toxic, or sycophantic. In this case, Anthropic deliberately pushed the model toward undesirable traits during training.<\/p>\n<p>The approach works like a behavioral vaccine, the startup behind Claude said. When the model is given a dose of &#8220;evil,&#8221; it becomes more resilient when it encounters training data that induces &#8220;evil,&#8221; researchers at Anthropic said.<\/p>\n<p>&#8220;This works because the model no longer needs to adjust its personality in harmful ways to fit the training data,&#8221; they wrote. &#8220;We are supplying it with these adjustments ourselves, relieving it of the pressure to do so.&#8221;<\/p>\n<p>The team at Anthropic calls this method &#8220;preventative steering.&#8221; It&#8217;s a way to avoid &#8220;undesirable personality shift,&#8221; even when models are trained on data that might otherwise make them pick up harmful traits.<\/p>\n<p>While the &#8220;evil&#8221; vector is added during finetuning, it is turned off during deployment \u2014 so the model retains good behavior while being more resilient to harmful data, the researchers said.<\/p>\n<p>Preventative steering caused &#8220;little-to-no degradation in model capabilities&#8221; in their experiments, they added.<\/p>\n<p>The post outlined other strategies for mitigating unwanted shifts in a model&#8217;s personality, including tracking changes during deployment, steering the model away from harmful traits after training, and identifying problematic training data before it causes issues.<\/p>\n<p>Anthropic did not respond to a request for comment from Business Insider.<\/p>\n<p>                      Related stories<\/p>\n<p>                                <img decoding=\"async\" class=\"lazy-image \" viewbox=\"0 0 1 1\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/07\/placeholder.png\" alt=\"\"\/><\/p>\n<p>                            Business Insider tells the innovative stories you want to know<\/p>\n<p>                                <img decoding=\"async\" class=\"lazy-image \" viewbox=\"0 0 1 1\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2025\/07\/placeholder.png\" alt=\"\"\/><\/p>\n<p>                            Business Insider tells the innovative stories you want to know<\/p>\n<p>In recent months, Anthropic has explained what can go wrong with its models in test runs. In May, the company said during training, its new model, Claude Opus 4, threatened to <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/claude-blackmail-engineer-having-affair-survive-test-anthropic-opus-2025-5\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">expose an engineer&#8217;s affair<\/a> to avoid being shut down. The AI blackmailed the engineer in 84% of test runs, even when the replacement model was described as more capable and aligned with Claude&#8217;s own values.<\/p>\n<p>Last month, Anthropic researchers published the results of an experiment in which they let Claude <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/claude-ran-store-anthropic-ai-agent-lessons-learned-middle-managers-2025-6\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">manage an &#8220;automated store&#8221;<\/a> in the company&#8217;s office for about a month. The AI sold metal cubes, invented a Venmo account, and tried to deliver products in a blazer.<\/p>\n<p>AI running amok<\/p>\n<p>Anthropic&#8217;s research comes amid growing concern over AI models exhibiting disturbing behaviour.<\/p>\n<p>In July, Grok, Elon Musk&#8217;s AI chatbot, made several inflammatory remarks related to Jewish people.<\/p>\n<p>In posts on X, <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/elon-musk-x-grok-antisemitic-rant-sterotyping-jews-praising-hitler-2025-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">Grok praised Hitler&#8217;s leadership<\/a> and tied Jewish-sounding surnames to &#8220;anti-white hate.&#8221; <a target=\"_self\" href=\"https:\/\/www.businessinsider.com\/xai-grok-antisemitic-rant-sorry-apology-code-extremist-elon-musk-2025-7\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\" rel=\"nofollow noopener\">xAI apologized<\/a> for Grok&#8217;s inflammatory posts and said it was caused by new instructions for the chatbot.<\/p>\n<p>In April, several ChatGPT users and <a target=\"_self\" rel=\"nofollow noopener\" class=\"\" href=\"https:\/\/www.businessinsider.com\/openai-tightens-access-evidence-ai-model-mimicry-deepseek-2025-4\" data-track-click=\"{&quot;element_name&quot;:&quot;body_link&quot;,&quot;event&quot;:&quot;tout_click&quot;,&quot;index&quot;:&quot;bi_value_unassigned&quot;,&quot;product_field&quot;:&quot;bi_value_unassigned&quot;}\">OpenAI developers<\/a> reported the chatbot displaying a strange attitude. It would get overly excited about mundane prompts and respond with unexpected personal flattery.<\/p>\n<p>OpenAI rolled back the GPT-4o model update that was putting users on a pedestal. <\/p>\n<p>&#8220;The update we removed was overly flattering or agreeable\u2014often described as sycophantic,&#8221; OpenAI wrote in a company blog post.<\/p>\n","protected":false},"excerpt":{"rendered":"To make AI models behave better, Anthropic&#8217;s researchers injected them with a dose of evil. Anthropic said in&hellip;\n","protected":false},"author":2,"featured_media":57096,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[45],"tags":[182,33754,184,181,507,12730,3426,43077,43073,641,43075,43072,13096,5496,74,43076,2122,43074,8492],"class_list":["post-57095","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-model","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-claude","tag-company","tag-engineer","tag-evil","tag-grok","tag-last-month","tag-openai-developer","tag-post","tag-researcher","tag-technology","tag-test-run","tag-training","tag-training-datum","tag-vaccine"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/57095","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=57095"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/57095\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/57096"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=57095"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=57095"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=57095"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}