{"id":675675,"date":"2026-05-16T23:59:40","date_gmt":"2026-05-16T23:59:40","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/675675\/"},"modified":"2026-05-16T23:59:40","modified_gmt":"2026-05-16T23:59:40","slug":"aws-found-bugs-in-60-of-software-requirements-its-fix-isnt-more-ai-its-a-50-year-old-logic-engine","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/675675\/","title":{"rendered":"AWS found bugs in 60% of software requirements. Its fix isn&#8217;t more AI \u2014 it&#8217;s a 50-year-old logic engine."},"content":{"rendered":"<p>The most expensive bugs in software aren\u2019t in the code. They\u2019re in the requirements that guide the code\u2019s construction.<\/p>\n<p>That\u2019s what AWS is trying to eliminate with new features in its <a href=\"https:\/\/thenewstack.io\/aws-kiro-brings-automated-reasoning-to-agentic-development\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Kiro<\/a> <a href=\"https:\/\/thenewstack.io\/antigravity-is-googles-new-agentic-development-platform\/\" type=\"link\" id=\"https:\/\/thenewstack.io\/antigravity-is-googles-new-agentic-development-platform\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">agentic development platform<\/a>. The primary new feature is <a href=\"https:\/\/kiro.dev\/blog\/deep-spec-analysis\/\" type=\"link\" id=\"https:\/\/kiro.dev\/blog\/deep-spec-analysis\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Requirements Analysis<\/a>, which tackles requirement bugs. These are contradictions, ambiguities, and gaps in specifications that get baked into design and code before anyone catches them. By the time they surface in production, tracing them back to a misread requirement can mean weeks of debugging.<\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/camikemiller\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Mike Miller<\/a>, director of AI product management at AWS, tells The New Stack that \u201ca bug in a requirement could be things that are contradicting requirements that imply two different things, ambiguities or gaps where a requirement might mean one thing to one developer but something slightly different to another.\u201d<\/p>\n<p>\u201cAnd so down the path of implementation, code testing, and then in production, maybe something doesn\u2019t work as expected, and you start rewinding,\u201d says Miller, who leads the Requirements Analysis initiative.<\/p>\n<p>The feature works in three stages, Miller explains. First, an LLM rewrites vague, natural-language requirements into precise, testable criteria. Second, that output gets translated into formal mathematical logic \u2014 what AWS calls a \u201cformal representation.\u201d Third, an <a href=\"https:\/\/en.wikipedia.org\/wiki\/Satisfiability_modulo_theories\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">SMT (satisfiability modulo theories) solver<\/a>, a type of automated reasoning engine, runs proofs against that logic to identify contradictions, ambiguities, undefined behaviors, and gaps. Findings surface to the developer as plain-language, two-option questions that Miller says can be resolved in about 10 to 15 seconds each.<\/p>\n<p>The term AWS keeps reaching for is prove. This is not the LLM flagging a probable issue \u2014 it is a formal reasoning engine demonstrating that no possible implementation can simultaneously satisfy two conflicting rules, the company says.<\/p>\n<p>\u201cAutomated reasoning allows us to take those requirements, look at them, identify gaps and ambiguities, and kind of address them up front,\u201d Miller says. \u201cThe LLM side does what it does best, and automated reasoning does what it does best.\u201d<\/p>\n<p><a href=\"https:\/\/www.linkedin.com\/in\/jasontandersen\/\" type=\"link\" id=\"https:\/\/www.linkedin.com\/in\/jasontandersen\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Jason Andersen<\/a>, an analyst with Moor Insights &amp; Strategy, tells The New Stack that \u201cAWS has been a pioneer in the idea\u00a0that LLM model correctness can be evaluated using diverse\u00a0algorithmic\u00a0models to improve accuracy.\u201d<\/p>\n<p>\u201cIt started with the use of Automated Reasoning in access control products such as\u00a0IAM,\u201d Andersen continues. \u201cThat success has started to spread into other AWS product lines. This is not the only method for judging LLM outputs. The more typical approach\u00a0is\u00a0to use\u00a0additional LLMs to inspect the\u00a0outputs and determine whether they make\u00a0sense.\u201d<\/p>\n<p>The neurosymbolic positioning<\/p>\n<p>The term <a href=\"https:\/\/thenewstack.io\/allegrograph-8-0-incorporates-neuro-symbolic-ai-a-pathway-to-agi\/\" type=\"link\" id=\"https:\/\/thenewstack.io\/allegrograph-8-0-incorporates-neuro-symbolic-ai-a-pathway-to-agi\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">neurosymbolic AI<\/a> refers to the combination of neural networks \u2014 the statistical, pattern-matching machinery behind LLMs \u2014 with symbolic logic, the rule-based, mathematically rigorous branch of AI that has been used for decades in formal verification and model checking, Miller says.<\/p>\n<p>\u201cSpeed without correctness just means you write wrong software faster.\u201d<\/p>\n<p>He uses the <a href=\"https:\/\/en.wikipedia.org\/wiki\/Pythagorean_theorem\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Pythagorean theorem<\/a> as an analogy to explain the difference in approach. An LLM trained on thousands of right triangles might infer the relationship between the sides and the hypotenuse. But it is inferring. It could be wrong. An automated reasoning system, by contrast, uses mathematical symbols to prove the relationship holds across every possible right triangle \u2014 not as a probability, but as a certainty, Miller says. <\/p>\n<p>Formal verification techniques built on this kind of symbolic logic have been used in hardware design and safety-critical software since the 1970s \u2014 some 50 years before the advent of LLMs. <\/p>\n<p>\u201cIt\u2019s not just about velocity,\u201d he notes. \u201cSpeed without correctness just means you write wrong software faster.\u201d<\/p>\n<p>Kiro was built around <a href=\"https:\/\/thenewstack.io\/vibe-coding-spec-driven\/\" type=\"link\" id=\"https:\/\/thenewstack.io\/vibe-coding-spec-driven\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">spec-driven development<\/a> from the start, tracing every line of generated code back to a documented requirement, Miller says. Requirements Analysis is meant to make that trace not just documented, but logically sound.<\/p>\n<p>In internal testing across 35 Kiro projects with more than 1,400 acceptance criteria, roughly 60% of first-draft requirements needed refinement before they could be reliably implemented, Miller says. But he said that is to be expected, as a first draft is a starting point.<\/p>\n<p>Why now<\/p>\n<p>AWS has been doing automated reasoning work quietly for years. The technology already appears in <a href=\"https:\/\/aws.amazon.com\/bedrock\/guardrails\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Bedrock Guardrails<\/a>, where a similar formal-logic pipeline can encode a chatbot\u2019s behavioral policy and validate responses against it mathematically, Miller says. It also appears in the\u00a0<a href=\"https:\/\/aws.amazon.com\/about-aws\/whats-new\/2026\/03\/policy-amazon-bedrock-agentcore-generally-available\/\" target=\"_blank\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\">Bedrock AgentCore policy<\/a>, which uses the same reasoning engine to determine when agents can use which tools under which circumstances.<\/p>\n<p>Requirements Analysis represents the first time that capability has been embedded directly in the development workflow, at the moment when specs are being written, Miller claims.<\/p>\n<p>\u201cWe are not seeing many evaluations applied at this point in the dev toolchain, let alone with a more advanced algorithmic technique.\u201d<\/p>\n<p>\u201cMy findings with Kiro are\u00a0that they have been very successful in pushing the envelope of features and getting to market first. In this\u00a0case,\u201d Andersen says. \u201cI would agree that they are ahead with this level of requirements reviews. We are not seeing\u00a0many evaluations applied at this point in the dev toolchain, let alone with a more advanced algorithmic technique.\u201d<\/p>\n<p>AWS has found that healthcare, finance, and other sectors where correctness is non-negotiable have been drawn to its automated reasoning capabilities, specifically because they need AI that doesn\u2019t hallucinate in sensitive contexts. The same pattern, AWS says, is emerging with agentic coding tools.<\/p>\n<p>In addition to Requirements Analysis, other new Kiro features include: Parallel Task Execution, which runs independent coding tasks concurrently to cut implementation time for large specs by roughly 75%, and Quick Plan, which generates a full set of requirements, design specs, and task breakdowns in a single pass after asking clarifying questions upfront.<\/p>\n<p>Kiro competes in a market with several popular AI coding tools, including Cursor, Codex, <a href=\"https:\/\/thenewstack.io\/claude-code-can-now-do-your-job-overnight\/\" type=\"link\" id=\"https:\/\/thenewstack.io\/claude-code-can-now-do-your-job-overnight\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Claude Code<\/a>, <a href=\"https:\/\/thenewstack.io\/github-copilot-a-powerful-controversial-autocomplete-for-developers\/\" type=\"link\" id=\"https:\/\/thenewstack.io\/github-copilot-a-powerful-controversial-autocomplete-for-developers\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">GitHub Copilot<\/a>, and Windsurf, among others.<\/p>\n<p>However, AWS says Kiro is widely used in the industry.<\/p>\n<p>Kiro\u2019s broader customer base already spans industries where getting it right matters as much as getting it done. <a href=\"https:\/\/www.socure.com\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Socure<\/a>, a digital identity verification and fraud prevention company, used Kiro\u2019s spec-driven development to complete a <a href=\"https:\/\/thenewstack.io\/scala-creator-proposes-lean-scala-for-simpler-code\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Scala<\/a>-to-<a href=\"https:\/\/thenewstack.io\/go\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Go<\/a> migration in two days. The project was originally scoped at three weeks. <\/p>\n<p><a href=\"https:\/\/www.nymbus.com\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Nymbus<\/a>, a banking technology provider, generates 80% of its <a href=\"https:\/\/thenewstack.io\/terraform-competitor-formae-expands-to-more-clouds\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Terraform<\/a> code, unit tests, and <a href=\"https:\/\/thenewstack.io\/a-practical-guide-to-data-driven-tests-with-playwright\/\" class=\"local-link\" rel=\"nofollow noopener\" target=\"_blank\">Playwright<\/a> object models with Kiro, cutting testing time on one project from 32 weeks to 7. Delta Air Lines reached its pilot program goals two quarters ahead of schedule. Nielsen saw a 25% increase in test coverage and a 40% decrease in time spent on documentation. Hughes Network Systems says Kiro specs eliminate the need to repeatedly re-establish context throughout the development workflow.<\/p>\n<p>The Kiro adoption list also includes Siemens, Rackspace Technology, Mondelez International, Appian, and Ericsson, alongside Amazon\u2019s own internal teams \u2014 Alexa+, Prime Video, Amazon Stores, and Fire TV among them.<\/p>\n<p>The leadership signal<\/p>\n<p>In addition to the Kiro feature launch, AWS announced that <a href=\"https:\/\/www.linkedin.com\/in\/shawn-bice-9205423\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Shawn Bice<\/a> has joined the company as VP of AI Services within Agentic AI, reporting to <a href=\"https:\/\/www.linkedin.com\/in\/swaminathansivasubramanian\/\" class=\"ext-link\" rel=\"external  nofollow noopener\" onclick=\"this.target=&#039;_blank&#039;;\" target=\"_blank\">Swami Sivasubramanian<\/a>, VP of Agentic AI at AWS. Bice will lead AWS\u2019s Automated Reasoning Group.<\/p>\n<p>In an internal memo to employees, Sivasubramanian wrote: \u201cWe are at an inflection point with Agentic AI, and I can\u2019t stress enough how critical AI and Automated Reasoning need to come together to build reliable and trustworthy agents.\u201d<\/p>\n<p>\u201cTo me, whether it\u2019s a better or more precise method is not the question,\u201d Andersen says.\u00a0\u201cMy question is: what\u2019s the impact on the human-in-the-loop? If AWS is better at locating an issue, that\u2019s a good thing, but ultimately it\u2019s going to come back to the developer to figure out what to do at this point. At some future point when we are automating more of the toolchain, any improvement of this type could be very valuable.\u201d<\/p>\n<p>AWS is betting that the next competitive axis in AI-assisted development is not how fast you can generate code, but how much you can trust what gets generated. Requirements Analysis is key to that bet.<\/p>\n<p>\t<a class=\"row youtube-subscribe-block\" href=\"https:\/\/youtube.com\/thenewstack?sub_confirmation=1\" target=\"_blank\" rel=\"nofollow noopener\"><\/p>\n<p>\n\t\t\t\tYOUTUBE.COM\/THENEWSTACK\n\t\t\t<\/p>\n<p>\n\t\t\t\tTech moves fast, don&#8217;t miss an episode. Subscribe to our YouTube<br \/>\n\t\t\t\tchannel to stream all our podcasts, interviews, demos, and more.\n\t\t\t<\/p>\n<p>\t\t\t\tSUBSCRIBE<\/p>\n<p>\t<\/a><\/p>\n<p>    Group<br \/>\n    Created with Sketch.<\/p>\n<p>\t\t<a href=\"https:\/\/thenewstack.io\/author\/darryl-taft\/\" class=\"author-more-link\" rel=\"nofollow noopener\" target=\"_blank\"><\/p>\n<p>\t\t\t\t\t<img decoding=\"async\" class=\"post-author-avatar\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2025\/07\/a95bb5bc-image-576x600.png\"\/><\/p>\n<p>\n\t\t\t\t\t\t\tDarryl K. Taft covers DevOps, software development tools and developer-related issues from his office in the Baltimore area. He has more than 25 years of experience in the business and is always looking for the next scoop. He has worked&#8230;\t\t\t\t\t\t<\/p>\n<p>\t\t\t\t\t\tRead more from Darryl K. Taft\t\t\t\t\t\t<\/p>\n<p>\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"The most expensive bugs in software aren\u2019t in the code. They\u2019re in the requirements that guide the code\u2019s&hellip;\n","protected":false},"author":2,"featured_media":675676,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[256,254,255,64,63,105],"class_list":["post-675675","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-au","tag-australia","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/675675","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=675675"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/675675\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/675676"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=675675"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=675675"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=675675"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}