{"id":805467,"date":"2026-08-07T07:06:08","date_gmt":"2026-08-07T07:06:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/805467\/"},"modified":"2026-08-07T07:06:08","modified_gmt":"2026-08-07T07:06:08","slug":"improving-fable-5-safeguards-anthropic","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/805467\/","title":{"rendered":"Improving Fable 5 Safeguards \\ Anthropic"},"content":{"rendered":"<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We\u2019re making updates to Claude Fable 5\u2019s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer \u201cfallbacks\u201d\u2014where the system switches to a less capable model after they make a biology-related query. In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces.1<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Fable 5 will thus be able to assist with a wider range of biology tasks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In practice, users should see far fewer fallbacks on everyday health and educational questions\u2014for example, interpreting lab results, understanding symptoms, and learning about biology in an educational context. Healthcare professionals will be able to receive more support from Fable 5 on clinical tasks.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We believe the greatest opportunity for AI to positively affect the world is in biology and medicine, and we&#8217;re investing significantly in building a responsible way to give biologists frontier access. Today, Fable still falls back to Opus 5 for requests we consider dual-use\u2014including virology, toxicology, and molecular design\u2014so it isn&#8217;t yet usable for professional biology research and drug development. We&#8217;re committed to closing that gap through trusted access pathways for frontier biology capabilities.<\/p>\n<p>Why we built strong biology safeguards<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Our objective is to get Fable 5\u2019s frontier capabilities into the hands of as many of our users as possible, as quickly as possible. However, to do so, we need to manage the increasing risks that come with models this capable. One such risk is in the field of biology: Fable 5 can now outperform experts on some highly complex biological tasks and provide operational support on others. That means that it can provide genuine assistance to a researcher developing a new medical treatment (which is the reason we\u2019re so keen to widen access to the model via both classifier improvements and trusted access programs). But in the wrong hands, those same capabilities could be used by a malicious actor, for example in developing a biological weapon. Our <a href=\"https:\/\/www-cdn.anthropic.com\/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf\" rel=\"nofollow noopener\" target=\"_blank\">capability assessments<\/a> show that Fable 5 could provide significant uplift to such an actor\u2014that is, it could provide them with capabilities they could not find anywhere else.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">It\u2019s often difficult to tell apart beneficial and harmful uses of AI in biology. For example, in some cases researching a treatment for a disease requires scientists to produce the dangerous compounds that cause that disease in the first place. This is most obvious for live vaccines, which require scientists to grow the same pathogen they\u2019re aiming to prevent. It\u2019s also the case for some medicines. To develop the drug captopril, which treats hypertension, scientists isolated toxic components of snake venom that crash blood pressure in humans. As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Sophisticated actors who wish to use our models to do harm know how to exploit this ambiguity to obscure their intent, making dangerous tasks look like ordinary research pursuits. The US Intelligence Community\u2019s <a href=\"https:\/\/www.dni.gov\/files\/ODNI\/documents\/assessments\/ATA-2026-Unclassified-Report.pdf\" rel=\"nofollow noopener\" target=\"_blank\">2026 Annual Threat Assessment<\/a> makes clear that such actors exist, and that advances in biotechnology including synthetic biology and genomic editing \u201ccould lead to novel biological threats.\u201d It notes that several state actors likely maintain active offensive biological and chemical weapons programs\u2014programs that could be accelerated by access to the raw capabilities of frontier AI models.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Because of our concerns about these \u201cdual-use\u201d capabilities (those that could be used for beneficial or harmful purposes, and where the line between them is not always easy to draw), we intentionally launched Fable 5 with almost all biology queries blocked. This enabled us to make the model available for users in other domains. We knew this would be frustrating for legitimate biology users: it would result in a high number of false positives in the near term, where users asking biology-related questions would have their requests blocked and sent to a less capable model. Nevertheless, we chose to make this tradeoff because the cost of Fable being misused in a dual-use domain like biology could potentially be <a href=\"https:\/\/cdn.sanity.io\/files\/4zrzovbb\/website\/0bacdc8440ea96e62a8766d99ebe1d4eea6d5f3a.pdf\" rel=\"nofollow noopener\" target=\"_blank\">catastrophic<\/a>.<\/p>\n<p>How our biology safeguards work<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">One of the core ways we protect against misuse in biology is via safety classifiers: smaller, automated AI systems that detect when Fable 5 is asked to perform a safeguarded biology task, or produce a harmful output (we&#8217;ve <a href=\"https:\/\/www.anthropic.com\/news\/fable-safeguards-jailbreak-framework\" rel=\"nofollow noopener\" target=\"_blank\">previously written<\/a> about our similar classifiers in the domain of cybersecurity).<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In the case of Fable 5, when a classifier fires, the model re-routes the user\u2019s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Developing precise, robust classifiers is not a straightforward task. For a classifier to work rapidly and consistently, it has to learn the difference between what we consider \u201cin scope\u201d and \u201cout of scope\u201d for the topics and queries we consider to be potentially harmful. It takes time and iteration to tune the classifiers, avoiding both false positives (where classifiers fire on out-of-scope content) and false negatives (where in-scope content is missed). We also require our classifiers to be robust to attempts to bypass them (known as jailbreaks), which requires even further research and testing.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Starting with a very broad biology classifier meant that we could give our users access to Fable 5 while we continued our research aimed at refining it. The alternative\u2014holding back the model until much more safeguards progress was made\u2014would have delayed the model\u2019s general access, and its potential benefits to our users, by weeks or months.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Over the past several weeks, we&#8217;ve carefully rewritten the classifier\u2019s constitution (which consists of a collection of rules to help the model discern between safeguarded and allowed content), taking care to carve out benign uses in detail. We solicited feedback on the changes from a diverse range of experts (both internal and external to Anthropic). We then developed updated training data for the classifier based on that constitution, and retrained it, and verified the new classifier would still generally trigger for harmful and dual-use research biology content but would now enable a wider range of benign and beneficial uses.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">As is illustrated in the diagram below, these updates meant that\u2014compared to at the time of Fable 5\u2019s launch\u2014the classifier will trigger for many fewer benign biology-related requests.<\/p>\n<p><img loading=\"lazy\" width=\"3840\" height=\"1855\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\"  src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2026\/08\/1786086368_918_image.webp\"\/>Illustration of our biology classifiers. Content that falls on the left-hand side of the classifier boundary is allowed; content that falls on the right-hand side is safeguarded (and is therefore blocked and sent instead to a less capable model);. Clearly harmful content (red), and content that is dual-use (orange), triggers the classifier and is blocked. We include a safety margin that includes content that is very likely benign but which is still blocked out of an abundance of caution (light green). Clearly benign content is in darker green.Upon its launch, Fable 5 had very broad classifiers (A) that triggered on a wide range of requests\u2014even ones that were almost certainly benign (those in the safety margin). The classifier boundary is thus very far to the left-hand side of the diagram. The update we are announcing today (B) means that many more benign requests are allowed by the classifier, which has become better at discerning subtle differences between benign and dual-use queries. The classifier boundary in the diagram has therefore moved further to the right-hand side.Conclusions<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">There\u2019s still much more to be done to refine our safeguards. There will inevitably remain false positives\u2014requests that fall within the classifier\u2019s safety margin where the request is very low-risk but where the classifier still fires. As we noted above, Fable will continue to block dual-use professional biology and drug development queries because of potential dual-use risk. We are fully committed to developing a safe, scalable path for researchers to use our most capable models via trusted access pathways.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We hope you\u2019ll continue to share your feedback with us so we can improve our safeguards even further.<\/p>\n","protected":false},"excerpt":{"rendered":"We\u2019re making updates to Claude Fable 5\u2019s biology safeguards in a way that substantially reduces false positives. Fable&hellip;\n","protected":false},"author":2,"featured_media":805468,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[45],"tags":[182,181,507,74],"class_list":["post-805467","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/805467","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=805467"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/805467\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/805468"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=805467"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=805467"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=805467"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}