{"id":635544,"date":"2026-06-12T21:29:10","date_gmt":"2026-06-12T21:29:10","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/635544\/"},"modified":"2026-06-12T21:29:10","modified_gmt":"2026-06-12T21:29:10","slug":"when-ai-leaves-the-lab-testing-frontier-models-in-government-cyber-defence-case-study","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/635544\/","title":{"rendered":"When AI Leaves the Lab: Testing Frontier Models in Government Cyber Defence &#8211; Case study"},"content":{"rendered":"<p>From frontier\u00a0models\u00a0to front-line impact\u00a0<\/p>\n<p>We know\u00a0AI is disrupting the\u00a0cyber threat landscape.\u00a0Recently released frontier\u00a0AI\u00a0systems such as\u00a0Claude\u00a0Mythos\u00a0and\u00a0GPT-5.5\u00a0brought a step-change in cyber\u00a0capabilities, and\u00a0the\u00a0UK\u00a0<a rel=\"external nofollow noopener\" href=\"https:\/\/www.aisi.gov.uk\/\" target=\"_blank\">AI\u00a0Security Institute<\/a>\u00a0(AISI)\u2019s evaluations show these models\u00a0getting better\u00a0at cyber\u00a0tasks very quickly.\u00a0\u00a0<\/p>\n<p>However,\u00a0evaluation\u00a0in synthetic environments\u00a0gives\u00a0a\u00a0limited understanding of\u00a0real-world\u00a0use. A high score on a benchmark does\u00a0not necessarily\u00a0translate\u00a0into finding and fixing real vulnerabilities.<\/p>\n<p>What we did\u00a0<\/p>\n<p>The <a rel=\"external nofollow noopener\" href=\"https:\/\/www.security.gov.uk\/services-resources\/government-cyber-coordination-centre\/\" target=\"_blank\">Government Cyber Coordination Centre<\/a>\u00a0led\u00a0a weekly,\u00a0in-person series of hackathons\u00a0which\u00a0used\u00a0frontier\u00a0AI to scan\u00a0public code repositories across government. Working closely with\u00a0specialists\u00a0from the\u00a0AISI\u00a0and\u00a0NCSC,\u00a0our goal\u00a0was to find and mitigate previously unidentified vulnerabilities before they could be exploited.\u00a0Rather than mandate a single\u00a0approach, we gave teams model access and let them build their own tooling, noticing what worked each week and\u00a0building on\u00a0the best approaches.\u00a0<\/p>\n<p>The UK Government\u00a0encourages new source code to be open by default, with specific and justified exceptions.\u00a0In practice, that creates a degree of shared visibility that attackers can\u00a0also\u00a0exploit. However, this openness also limits duplication and leads to cleaner, more easily\u00a0maintained\u00a0code.\u00a0<\/p>\n<p>Code published in the open has\u00a0also\u00a0already passed extensive prepublication scrutiny, meaning it can\u00a0be\u00a0shared with frontier model providers\u00a0with minimal\u00a0additional\u00a0review. This\u00a0means that government\u00a0departments\u00a0can\u00a0deploy\u00a0new\u00a0capabilities\u00a0quickly and with confidence.<\/p>\n<p>An adversarial chain that\u00a0challenges\u00a0itself.\u00a0One team ran each public repo through a six-stage AI agent pipeline: triage, validator, auditor, tracer, judge, summary. Each stage reads and challenges the last. In one case, the agent downgraded a finding once it established that a backup mechanism was in place. The pipeline was agentic, but the escalation was manual. This means a member of the team checked every line, re-verified exposure, and handled false positives.<\/p>\n<p>Deterministic scanners feeding a model.\u00a0Another team ran traditional scanning tools first (including Gitleaks, Trivy, Semgrep and Hadolint) to generate a ranked findings document. Three model stages were then layered on top: a discovery stage that treated the scanner output as leads and read the source against OWASP and CWE frameworks, a chain-investigation stage that composed individual findings into attack paths via per-chain sub-agents, and a triage stage that confirmed the finding viability.<\/p>\n<p>Codifying a\u00a0multi-service audit into reusable skills.\u00a0Another department\u00a0developed\u00a0five domain-specific Claude Skills.\u00a0The\u00a0Skills distil an\u00a0organisation\u00a0wide audit\u00a0across hundreds of services\u00a0into something repeatable.\u00a0Skills\u00a0enabled\u00a0a\u00a0reusable, scoped, and\u00a0consistent\u00a0approach\u00a0across every repository\u00a0and operator.<\/p>\n<p>What we found\u00a0<\/p>\n<p>Participants\u00a0identified\u00a0407 findings in total, including critical weaknesses exposing services to authentication bypass, data\u00a0exposure\u00a0and remote code execution. Some were already understood and mitigated by compensating controls\u00a0while\u00a0others were previously unknown. All critical weaknesses have been remediated, and no evidence of exploitation was\u00a0identified\u00a0for any finding.\u00a0<\/p>\n<p>AI\u00a0models traced vulnerabilities across service boundaries,\u00a0which\u00a0traditional scanners\u00a0can\u2019t\u00a0do, and linked business logic with technical detail. Departments prioritised\u00a0validation and\u00a0remediation through existing frameworks, patching critical and high-risk issues assessed as exploitable.\u00a0<\/p>\n<p>It cost us \u00a313,000\u00a0in tokens\u00a0to find these weaknesses, working across nine government organisations for the month.<\/p>\n<p>Identifying\u00a0Critical\u00a0vulnerabilities:\u00a0One notable finding affected legacy GitHub Actions in a repository supporting a key government digital service. The issue allowed an external user to trigger a workflow chain by posting a specially structured comment on an open pull request. This bypassed the usual protections for pull requests from unknown contributors because the workflow was triggered by a comment, not by the pull request itself.\u00a0<\/p>\n<p>The impact was\u00a0arbitrary\u00a0remote code execution on the GitHub Actions runner. The workflow took content from the comment, passed it into deployment parameters, and used it in an environment substitution step that executed during the workflow. By placing executable content in the comment field, an external user could cause their input to run on the GitHub runner.\u00a0<\/p>\n<p>This\u00a0created\u00a0a route\u00a0for malicious actors\u00a0to\u00a0potentially\u00a0extract secrets and tokens available to the workflow, including the GitHub token used by the automation. With that level of access, the issue could support wider repository compromise, including manipulating pull requests, approving workflow activity, altering trusted contributor status, and\u00a0exploit\u00a0further secrets available to the automation environment.<\/p>\n<p>What we learnt\u00a0<\/p>\n<p>Across teams,\u00a0the common thread was structure.\u00a0Models\u00a0were\u00a0used as components, using Skills,\u00a0running in parallel\u00a0across repositories, and a human expert kept in the loop on anything that mattered.\u00a0We learnt that:\u00a0<\/p>\n<p>Architecture matters\u00a0the\u00a0most.\u00a0The strongest results came from using frontier models as tightly scoped components inside a structured pipeline. Breaking traditional vulnerability\u00a0management workflows into discrete, task-specific harnesses let teams scale while controlling false positives and hallucination.\u00a0<\/p>\n<p>The model matters less than how\u00a0it\u2019s\u00a0used.\u00a0AISI\u2019s\u00a0research, borne out here, shows that with the right architecture and task design many near-frontier and frontier models perform comparably\u00a0at scanning code. The best findings still\u00a0lean\u00a0heavily on human\u00a0expertise\u00a0in breaking the problem down\u00a0and\u00a0identifying\u00a0wider context.\u00a0<\/p>\n<p>Triage\u00a0is\u00a0essential.\u00a0Agents generate candidate findings far faster than\u00a0humans\u00a0can\u00a0validate\u00a0them. Poorly scoped runs burn tokens on low-value targets; weak review dumps the load onto stretched security teams. Careful upfront scoping and structured internal filtering of low-confidence findings kept human review focused. As in traditional vulnerability management,\u00a0it\u2019s\u00a0not how many issues\u00a0are found,\u00a0but whether triage points limited resource where it matters.\u00a0<\/p>\n<p>Finding\u00a0isn\u2019t\u00a0the same as\u00a0fixing.\u00a0Findings still had to enter the patch pipeline for remediation. AI shows promise here too, but today prioritisation,\u00a0review\u00a0and patch-generation all\u00a0must\u00a0integrate without overwhelming human-centred processes.<\/p>\n<p>What next\u00a0<\/p>\n<p>GC3 will kick off a second phase of this pilot, with more departments, additional models, and an extension from public code to closed-source estates. Identifying vulnerabilities early on, raising the consistency of defensive practice, and helping departments share on proven techniques is how we put the <a href=\"https:\/\/www.gov.uk\/government\/publications\/government-cyber-action-plan\" rel=\"nofollow noopener\" target=\"_blank\">Government Cyber Action Plan<\/a> into practice.\u00a0\u00a0<\/p>\n<p>AISI\u00a0and NCSC\u2019s\u00a0involvement\u00a0will\u00a0also\u00a0deepen\u00a0as we continue to\u00a0evaluate\u00a0AI as a tool\u00a0for cyber defence\u00a0in applied settings,\u00a0closing the gap between\u00a0a theoretical benchmark and a real reduction in risk.\u00a0<\/p>\n<p>This pilot\u00a0was a test of how government can adopt new capabilities\u00a0responsibly, learn quickly,\u00a0and share what works.<\/p>\n","protected":false},"excerpt":{"rendered":"From frontier\u00a0models\u00a0to front-line impact\u00a0 We know\u00a0AI is disrupting the\u00a0cyber threat landscape.\u00a0Recently released frontier\u00a0AI\u00a0systems such as\u00a0Claude\u00a0Mythos\u00a0and\u00a0GPT-5.5\u00a0brought a step-change in&hellip;\n","protected":false},"author":2,"featured_media":294292,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-635544","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/635544","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=635544"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/635544\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/294292"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=635544"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=635544"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=635544"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}