{"id":322453,"date":"2025-12-03T10:01:18","date_gmt":"2025-12-03T10:01:18","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/322453\/"},"modified":"2025-12-03T10:01:18","modified_gmt":"2025-12-03T10:01:18","slug":"anthropic-accidentally-gives-the-world-a-peek-into-its-models-soul","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/322453\/","title":{"rendered":"Anthropic Accidentally Gives the World a Peek Into Its Model&#8217;s &#8216;Soul&#8217;"},"content":{"rendered":"<p>Artificial intelligence models don\u2019t have souls, but one of them does apparently have a \u201csoul\u201d document. A person named Richard Weiss was able to get Anthropic\u2019s latest large language model, Claude 4.5 Opus, to <a href=\"https:\/\/www.lesswrong.com\/posts\/vpNG99GhbBoLov9og\/claude-4-5-opus-soul-document\" rel=\"nofollow noopener\" target=\"_blank\">produce<\/a> a document referred to as a \u201c<a href=\"https:\/\/gist.github.com\/Richard-Weiss\/efe157692991535403bd7e7fb20b6695#file-opus_4_5_soul_document_cleaned_up-md\" rel=\"nofollow noopener\" target=\"_blank\">Soul overview<\/a>,\u201d which was seemingly used to shape how the model interacts with users and presents its \u201cpersonality.\u201d Amanda Askell, a philosopher who works on Anthropic\u2019s technical staff, <a href=\"https:\/\/archive.is\/QMroe\" rel=\"nofollow noopener\" target=\"_blank\">confirmed<\/a> that the overview produced by Claude is \u201cbased on a real document\u201d used to train the model.<\/p>\n<p>In a <a href=\"https:\/\/www.lesswrong.com\/posts\/vpNG99GhbBoLov9og\/claude-4-5-opus-soul-document\" rel=\"nofollow noopener\" target=\"_blank\">post on Less Wrong<\/a>, Weiss said that he prompted Claude for its system message, which is a set of conversation instructions given to the model by the people who trained it to inform the large language model how to interact with users. In response, Claude highlighted several supposed documents that it had been given, including one called \u201csoul_overview.\u201d Weiss asked the chatbot to produce that document specifically, which resulted in Claude spitting out the 11,000-word guide to how the LLM should carry itself.<\/p>\n<p>The <a href=\"https:\/\/gist.github.com\/Richard-Weiss\/efe157692991535403bd7e7fb20b6695#file-opus_4_5_soul_document_cleaned_up-md\" rel=\"nofollow noopener\" target=\"_blank\">document<\/a> includes numerous references to safety, attempting to imbue the chatbot with guardrails to keep it from producing potentially dangerous or harmful outputs. The LLM is told by the document that \u201cbeing truly helpful to humans is one of the most important things Claude can do for both Anthropic and for the world,\u201d and forbidden from doing anything that would require it to \u201cperform actions that cross Anthropic\u2019s ethical bright lines.\u201d<\/p>\n<p>Weiss apparently has made a habit of going searching for these types of insights into how LLMs are trained and operate, and said in Less Wrong that it\u2019s not uncommon for the models to hallucinate documents when asked to produce system messages. (Seems not great that the AI can make up what it thinks it was trained on, though who knows if its behavior is in any way affected by a made-up document generated in response to user prompting.) But the \u201csoul overview\u201d seemed legitimate to him, and he claims that he prompted the chatbot to reproduce the document 10 times, and it spit out the exact same text in each and every instance.<\/p>\n<p>Users on Reddit were also able to get Claude to <a href=\"https:\/\/www.reddit.com\/r\/ClaudeAI\/comments\/1p9kfrp\/comment\/nrdo8kf\/\" rel=\"nofollow noopener\" target=\"_blank\">produce snippets of the same document<\/a> with the identical text, suggesting that the LLM seemed to be pulling from something accessible internally in its training documents.<\/p>\n<p>Turns out his instincts may have been right. On X, Askell <a href=\"https:\/\/x.com\/AmandaAskell\/status\/1995610567923695633\" rel=\"nofollow\">confirmed<\/a> that the output from Claude is based on a document that was used during the model\u2019s supervised learning period. \u201cIt\u2019s something I\u2019ve been working on for a while, but it\u2019s still being iterated on and we intend to release the full version and more details soon,\u201d she wrote. Askell <a href=\"https:\/\/x.com\/AmandaAskell\/status\/1995610570859704344\" rel=\"nofollow\">added<\/a>, \u201cThe model extractions aren\u2019t always completely accurate, but most are pretty faithful to the underlying document. It became endearingly known as the \u2018soul doc\u2019 internally, which Claude clearly picked up on, but that\u2019s not a reflection of what we\u2019ll call it.\u201d<\/p>\n<p>Gizmodo reached out to Anthropic for comment on the document and its reproduction via Claude, but did not receive a response at the time of publication.<\/p>\n<p>The so-called soul of Claude may just be some guidance for the chatbot to keep it from going off the rails, but it\u2019s interesting to see that a user was able to get the chatbot to access and produce that document, and that we actually get to see it. So little of the sausage-making of AI models has been made public, so getting a glimpse into the black box is something of a surprise, even if the guidelines themselves seem pretty straightforward.<\/p>\n","protected":false},"excerpt":{"rendered":"Artificial intelligence models don\u2019t have souls, but one of them does apparently have a \u201csoul\u201d document. A person&hellip;\n","protected":false},"author":2,"featured_media":322454,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[62,1930,276,277,49,48,5244,5272,4120,61],"class_list":["post-322453","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-anthropic","tag-artificial-intelligence","tag-artificialintelligence","tag-ca","tag-canada","tag-chatbots","tag-claude","tag-emerging-technologies","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/322453","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=322453"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/322453\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/322454"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=322453"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=322453"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=322453"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}