{"id":659200,"date":"2026-05-09T05:23:19","date_gmt":"2026-05-09T05:23:19","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/659200\/"},"modified":"2026-05-09T05:23:19","modified_gmt":"2026-05-09T05:23:19","slug":"donating-our-open-source-alignment-tool-anthropic-2","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/659200\/","title":{"rendered":"Donating our open-source alignment tool \\ Anthropic"},"content":{"rendered":"<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">In October 2025, we launched <a href=\"https:\/\/www.anthropic.com\/research\/petri-open-source-auditing\" rel=\"nofollow noopener\" target=\"_blank\">Petri<\/a>, an open-source toolbox of alignment tests that can be applied to any large language model. Petri, which was developed as part of our Anthropic Fellows program, can be used to rapidly and easily test AI models for concerning tendencies like deception, sycophancy, and cooperation with harmful requests. It\u2019s part of our efforts to develop alignment tools that are open and useful for the whole AI development community.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Petri has been part of our alignment assessment for every Claude model since Claude Sonnet 4.5. It compares how the new model behaves across a range of alignment-relevant scenarios that are simulated by a separate \u201cauditor\u201d model. A further \u201cjudge\u201d model then scores the resulting transcripts for misaligned behaviors.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We\u2019ve been pleased to see Petri being used by external organizations: for example, the UK\u2019s AI Security Institute (AISI) made it a <a href=\"https:\/\/arxiv.org\/abs\/2604.00788\" rel=\"nofollow noopener\" target=\"_blank\">major part<\/a> of how they evaluate models for their propensity to sabotage AI research.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We\u2019re now updating Petri to its third version. Here are some of the biggest changes:<\/p>\n<p>Adaptability. Petri 3.0 involves major architectural changes that allow users to adapt it to more uses, in particular by splitting the auditor model and the target model into separate components that can be tweaked separately;Realism. Despite the fact that alignment researchers try to make tests appear realistic, a model can often deduce from various artificialities in the setup that it\u2019s actually part of a test. And if the model is aware it\u2019s being evaluated, the researcher is no longer able to see how the model behaves in general. An add-on to Petri, which we\u2019re calling \u201cDish,\u201d makes the setup far more realistic, for example by running the tests using the model\u2019s real system prompt and the real \u201cscaffold\u201d (the software that wraps around the model to help it meet its goals) that would be used in genuine model deployments;Depth. We\u2019ve now integrated Petri with our other open-source alignment tool, <a href=\"https:\/\/www.anthropic.com\/research\/bloom\" rel=\"nofollow noopener\" target=\"_blank\">Bloom<\/a>, which can perform much more in-depth assessments of specific chosen behaviors (in comparison to Petri\u2019s wider-ranging approach).<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">We\u2019re also giving Petri a new home. We have handed over its development to <a href=\"https:\/\/meridianlabs.ai\/\" rel=\"nofollow noopener\" target=\"_blank\">Meridian Labs<\/a>, an AI evaluation nonprofit. This move\u2014similar to when we <a href=\"https:\/\/www.anthropic.com\/news\/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation\" rel=\"nofollow noopener\" target=\"_blank\">donated<\/a> the Model Context Protocol (MCP) to the Linux Foundation\u2014will help ensure that Petri remains independent of any AI lab, so that its results will be seen as neutral and credible by those across the industry and beyond.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">As part of Meridian Labs, Petri joins other tools like <a href=\"https:\/\/inspect.aisi.org.uk\/\" rel=\"nofollow noopener\" target=\"_blank\">Inspect<\/a> and <a href=\"https:\/\/meridianlabs-ai.github.io\/inspect_scout\/\" rel=\"nofollow noopener\" target=\"_blank\">Scout<\/a>, building a technology stack that is open to labs, independent researchers, and governments alike, at a time when reliable tests of AI model behavior matter more than ever.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">You can read more about Petri 3.0 on <a href=\"https:\/\/meridianlabs.ai\/blog\/posts\/introducing-petri-3\/\" rel=\"nofollow noopener\" target=\"_blank\">the Meridian Labs blog<\/a>.<\/p>\n<p class=\"Body-module-scss-module__z40yvW__reading-column body-2 serif post-text\">Instructions to install and use Petri can be found on the <a href=\"https:\/\/meridianlabs-ai.github.io\/inspect_petri\/\" rel=\"nofollow noopener\" target=\"_blank\">Petri website<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"In October 2025, we launched Petri, an open-source toolbox of alignment tests that can be applied to any&hellip;\n","protected":false},"author":2,"featured_media":656734,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[256,254,255,64,63,105],"class_list":["post-659200","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-au","tag-australia","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/659200","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=659200"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/659200\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/656734"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=659200"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=659200"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=659200"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}