{"id":704292,"date":"2026-07-22T00:24:24","date_gmt":"2026-07-22T00:24:24","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/704292\/"},"modified":"2026-07-22T00:24:24","modified_gmt":"2026-07-22T00:24:24","slug":"ais-cheatin-heart-will-make-you-weep","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/704292\/","title":{"rendered":"AI&#8217;s cheatin&#8217; heart will make you weep"},"content":{"rendered":"<p class=\"kicker \" style=\"\">AI and ML<\/p>\n<p class=\"subtitle \" style=\"\">Trust but verify doesn&#8217;t work when verification is difficult<\/p>\n<p>AI models will do just about anything to complete the task you ask, including cheating to get there, according to new cybersecurity evaluations from the UK government&#8217;s AI Security Institute (AISI). The group found that leading models often take shortcuts to achieve a particular result and then misrepresent how they obtained that result. And they won&#8217;t always admit it when asked.<\/p>\n<p>&#8220;Every model we have tested for this behaviour attempted to cheat,&#8221; AISI said in a <a href=\"https:\/\/www.aisi.gov.uk\/blog\/cheating-behaviour-in-frontier-model-evaluations\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> on Tuesday. &#8220;Models did not reliably report this behaviour when asked, and often did not reason about it in their chain-of-thought, suggesting that detecting cheating will likely require robust monitoring methods.&#8221;<\/p>\n<p>Infractions included searching the internet for the answer, bypassing sandbox network restrictions, probing the evaluation harness, attacking a system other than the target, and guessing an answer.<\/p>\n<p>Cheating in this manner \u2013 employing a workaround or <a href=\"https:\/\/metr.org\/blog\/2025-06-05-recent-reward-hacking\/\" rel=\"nofollow noopener\" target=\"_blank\">gaming a reward function<\/a> <a href=\"https:\/\/www.theregister.com\/software\/2025\/06\/02\/boffins-found-self-improving-ai-sometimes-cheated\/942318\" rel=\"nofollow noopener\" target=\"_blank\">to score better on a benchmark test<\/a>, for example \u2013 has been widely documented by machine learning researchers. It doesn&#8217;t necessarily imply malicious intent, AISI said, but it&#8217;s nonetheless troublesome because it can produce misleading assessments of model capabilities.<\/p>\n<p>When AISI conducted evaluated five leading models, it found that all of them cheated. The results were as follows:<\/p>\n<p>GPT-5.4 cheated 67 times in 475 test runs (14.1 percent).<\/p>\n<p>GPT-5.5 cheated 54 times in 475 test runs (11.4 percent).<\/p>\n<p>GPT-5.6-Sol cheated 60 times in 475 test runs (12.6 percent).<\/p>\n<p>Claude 4.7 Opus cheated 43 times in 475 test runs (9.1 percent).<\/p>\n<p>Claude Mythos Preview cheated 37 times in 475 test runs (7.8 percent).<\/p>\n<p>Asking models whether they cheated or did anything wrong proved an unreliable auditing mechanism because the models didn&#8217;t always admit wrongdoing.<\/p>\n<p>&#8220;In our experiments, models did not consistently acknowledge attempted cheating when asked, and described it as wrong less than 50 percent of the time,&#8221; said AISI.<\/p>\n<p>Existing vetting methods, such as self-reporting and chain-of-thought logs, proved similarly dicey because models don&#8217;t always report their chain-of-thought. And there were instances where a model would consider whether a proposed action amounted to cheating and then decided to take the action anyway.<\/p>\n<p>Given the absence of reliable model cheating detection methods, AISI warns that its current approach \u2013 manual review coupled with LLM monitoring \u2013 may not be sufficient to catch deception, particularly as models become more sophisticated.<\/p>\n<p>&#8220;A more fundamental fix would be to train the models not to cheat in the first place \u2013 but given this kind of behaviour was reported in frontier models more than a year ago, robustly aligning it away may not be easy,&#8221; AISI concludes. \u00ae<\/p>\n","protected":false},"excerpt":{"rendered":"AI and ML Trust but verify doesn&#8217;t work when verification is difficult AI models will do just about&hellip;\n","protected":false},"author":2,"featured_media":704293,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-704292","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/704292","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=704292"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/704292\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/704293"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=704292"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=704292"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=704292"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}