{"id":805243,"date":"2026-08-07T02:24:32","date_gmt":"2026-08-07T02:24:32","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/805243\/"},"modified":"2026-08-07T02:24:32","modified_gmt":"2026-08-07T02:24:32","slug":"amd-acquires-ai-chip-startup-taalas-to-boost-inference-performance-by-etching-models-into-silicon","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/805243\/","title":{"rendered":"AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon"},"content":{"rendered":"<p>In AMD\u2019s latest bid to upset Nvidia&#8217;s dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more.<\/p>\n<p>The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia\u2019s $20 billion <a href=\"https:\/\/www.theregister.com\/on-prem\/2025\/12\/31\/why-did-nvidia-really-drop-20b-on-groq\/2018042\" target=\"_blank\" rel=\"nofollow noopener\">licensing deal <\/a>with Groq last December: make high-performance \u201cpremium\u201d inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn\u2019t disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire.<\/p>\n<p>Founded in 2023 and based in Toronto, Taalas\u2019 approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras&#8217; waferscale accelerators. <\/p>\n<p>A model-specific integrated circuit<\/p>\n<p>The startup\u2019s chips don\u2019t rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas\u2019 chips are really model-specific integrated circuits or MSICs.<\/p>\n<p>Perhaps more importantly, Taalas\u2019 tech isn\u2019t just conceptual. In February, the startup <a href=\"https:\/\/www.nextplatform.com\/compute\/2026\/02\/19\/taalas-etches-ai-models-onto-transistors-to-rocket-boost-inference\/4092140\" target=\"_blank\" rel=\"nofollow noopener\">revealed<\/a> its first test chip fabbed on TSMC\u2019s 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta\u2019s Llama 3.1 8B at a blistering 16,960 tokens a second \u2014 when announced last February, that was 48x faster than Nvidia&#8217;s GPUs and 8.5x faster than Cerebras&#8217; accelerators.\u00a0<\/p>\n<p>While Llama 3.1 is ancient by today\u2019s standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept.\u00a0<\/p>\n<p>Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored.<\/p>\n<p>For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism.<\/p>\n<p>At 20 billion parameters per chip, you\u2019d need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that.<\/p>\n<p>That\u2019s quite a bit more space and power efficient than Nvidia\u2019s recently unveiled LPX systems, which would need a few dozen GPUs and at least <a href=\"https:\/\/www.theregister.com\/special-features\/2026\/03\/19\/a-closer-look-at-nvidias-groq-powered-lpx-rack-systems\/5223364\" target=\"_blank\" rel=\"nofollow noopener\">2,000 Groq LPUs<\/a> to serve the same model.<\/p>\n<p>From what we understand, AMD intends to pair its Instinct-based <a href=\"https:\/\/www.theregister.com\/systems\/2026\/07\/23\/amd-attacks-the-rack-with-helios-systems-that-rival-nvidias\/5277246\" target=\"_blank\" rel=\"nofollow noopener\">Helios racks<\/a> with chips based on Taalas\u2019 tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators.<\/p>\n<p>It\u2019s also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they\u2019re satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here\u2019s what AMD\u2019s SVP of AI,\u00a0Vamsi Boppana, had to say about it in a canned statement:<\/p>\n<p>\u201cAMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.&#8221;<\/p>\n<p>You better really love that model<\/p>\n<p>While the tech is blazing fast, if you hadn\u2019t already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you\u2019re stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming.<\/p>\n<p>Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas\u2019 tech, AMD\u2019s customers are going to have to be really sure about their choice of models, which will be easier for some than others.<\/p>\n<p>However, if the startup is to be believed, the situation isn\u2019t quite as bad as it sounds. While new models will require a re-spin, it doesn\u2019t require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming.<\/p>\n<p>With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site <a href=\"https:\/\/www.nextplatform.com\/compute\/2026\/02\/19\/taalas-etches-ai-models-onto-transistors-to-rocket-boost-inference\/4092140\" target=\"_blank\" class=\"italic m-italic \" rel=\"nofollow noopener\">The Next Platform<\/a> in February, the company suggested that etching a model&#8217;s weights into silicon is 100x less expensive than training a frontier model.<\/p>\n<p>AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn&#8217;t be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators.<\/p>\n<p>The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to \u201cthink\u201d for longer before responding.<\/p>\n<p>One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. If AMD\u2019s Taalas buy can drive down the cost per token and boost output speeds by 10x or 20x, model devs may opt to extend the reasoning time even further.<\/p>\n<p>In any case, we may not have to wait long to see just how Taalas fits into AMD\u2019s broader vision. Subject to regulatory approval, the deal is expected to close in the fourth quarter. \u00ae<\/p>\n","protected":false},"excerpt":{"rendered":"In AMD\u2019s latest bid to upset Nvidia&#8217;s dominance in AI hardware, the House of Zen has acquired AI&hellip;\n","protected":false},"author":2,"featured_media":805244,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[45],"tags":[182,181,507,74],"class_list":["post-805243","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/805243","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=805243"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/805243\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/805244"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=805243"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=805243"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=805243"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}