{"id":686845,"date":"2026-05-22T18:04:12","date_gmt":"2026-05-22T18:04:12","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/686845\/"},"modified":"2026-05-22T18:04:12","modified_gmt":"2026-05-22T18:04:12","slug":"ex-google-deepmind-researcher-warns-benchmarks-wont-save-us","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/686845\/","title":{"rendered":"Ex-Google DeepMind Researcher Warns Benchmarks Won&#8217;t Save Us"},"content":{"rendered":"<p>Remember when there was that stretch of time where people were leaving AI companies and every one of their farewell messages boiled down to, \u201cThis is going to kill us all?\u201d Lun Wang, a researcher at Google\u2019s DeepMind, <a href=\"https:\/\/x.com\/lunwang1996\/status\/2056222588054237329?s=20\" rel=\"nofollow\">recently announced<\/a> he was departing from the company and may have reignited the trend by warning that current benchmarking tests aren\u2019t capable of truly evaluating risks presented by evolving AI models.<\/p>\n<p><a href=\"https:\/\/x.com\/lunwang1996\/status\/2056222588054237329?s=20\" rel=\"nofollow\">On X<\/a>, Wang noted that before deciding to depart from DeepMind, he had been thinking a lot about how AI models are evaluated. \u201cWe\u2019re good at evaluating the models we have. We\u2019re much worse at evaluating the models we\u2019re about to build \u2014 especially if they cross into a new capability regime. We will have self-evolving models, but before that, we need self-evolving evaluations,\u201d he wrote.<\/p>\n<p>He expanded on the idea in a <a href=\"https:\/\/wanglun1996.github.io\/blog\/your-evals-will-break.html\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a>, in which he explained further: \u201cMost benchmarks, safety evals, and red-teaming protocols implicitly assume the next model is a stronger version of the current one. If it\u2019s a different kind of thing, our entire evaluation infrastructure breaks silently.\u201d Basically, if we\u2019re counting on the current methods of stress testing AI to catch malicious behavior that we haven\u2019t already considered, we\u2019re probably shit out of luck.<\/p>\n<p>What would that look like? Wang offered an example:<\/p>\n<p>\u201cImagine a model that, at some scale, develops the ability to strategically withhold information to achieve goals \u2014 not lying exactly, but selectively omitting facts in ways that steer conversations toward outcomes its training process accidentally reinforced. Your existing honesty benchmarks wouldn\u2019t catch this, because they test for factual accuracy, not for strategic omission. Your safety classifiers wouldn\u2019t flag it, because the individual outputs are all technically true.\u201d<\/p>\n<p>In that scenario, benchmarks and safety checks wouldn\u2019t even know what to look for. They would monitor the risks that they are designed to watch out for, while the more nefarious functions slip right by. That would be bad!<\/p>\n<p>Wang did offer a solution\u2026 kinda. Basically, build better evaluations\u2014ones that can evolve as models do. Sounds like a good idea, maybe someone who is still working at these companies could go ahead and get started on that.<\/p>\n<p>Wang isn\u2019t the first to raise an alarm about the risks surrounding poor benchmarking. The method of evaluation has frequently been criticized for <a href=\"https:\/\/www.theregister.com\/software\/2025\/11\/07\/ai-benchmarks-hampered-by-bad-science\/1106339\" rel=\"nofollow noopener\" target=\"_blank\">failing to meaningfully define what it aims to measure<\/a> and being too rigidly tied to singular evaluation goals that <a href=\"https:\/\/removepaywalls.com\/https:\/\/www.technologyreview.com\/2026\/03\/31\/1134833\/ai-benchmarks-are-broken-heres-what-we-need-instead\/\" rel=\"nofollow noopener\" target=\"_blank\">often don\u2019t even reflect the way models are actually used in real life<\/a>. Benchmarking has become the de facto measure of model success across the industry, which has also led to companies effectively gaming the system by <a href=\"https:\/\/www.mindstudio.ai\/blog\/benchmark-gaming-ai-inflated-scores-explained\" rel=\"nofollow noopener\" target=\"_blank\">training against the test<\/a> and inflating their scores.<\/p>\n<p>If there were a benchmark for being a good benchmark, it seems the current benchmarks would fail.<\/p>\n","protected":false},"excerpt":{"rendered":"Remember when there was that stretch of time where people were leaving AI companies and every one of&hellip;\n","protected":false},"author":2,"featured_media":21408,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[62,19089,276,277,214,49,48,32136,15740,61],"class_list":["post-686845","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-models","tag-artificial-intelligence","tag-artificialintelligence","tag-benchmarks","tag-ca","tag-canada","tag-deepmind","tag-google-deepmind","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/686845","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=686845"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/686845\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/21408"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=686845"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=686845"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=686845"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}