{"id":532245,"date":"2026-07-09T02:28:16","date_gmt":"2026-07-09T02:28:16","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/532245\/"},"modified":"2026-07-09T02:28:16","modified_gmt":"2026-07-09T02:28:16","slug":"intel-backed-ai-chip-startup-sambanova-breathes-new-life-into-aging-nvidia-gpus-in-latest-benchmarks","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/532245\/","title":{"rendered":"Intel-backed AI chip startup SambaNova breathes new life into aging Nvidia GPUs in latest benchmarks"},"content":{"rendered":"<p class=\"kicker \" style=\"\">ai and ml<\/p>\n<p class=\"subtitle \" style=\"\">Third-party testing shows heterogeneous compute platform combining H200s and SN50 RDUs churning out 763 tok\/s in MiniMax M2.7<\/p>\n<p>Intel&#8217;s big bet on SambaNova appears to be paying off in a big way. This week, the AI chip startup shared benchmark results showing its latest generation of AI acceleration, which combines Nvidia GPUs and the company&#8217;s accelerators, beating GPU-only inference platforms by a wide margin.<\/p>\n<p>The testing, conducted by the AI <a href=\"https:\/\/sambanova.ai\/blog\/sn50-runs-fastest-minimax-speeds-in-the-world\" rel=\"nofollow noopener\" target=\"_blank\">benchmarking<\/a> gurus at Artificial Analysis, showed SambaNova&#8217;s SN50-series accelerators, <a href=\"https:\/\/www.theregister.com\/on-prem\/2026\/02\/24\/sambanova-raises-350m-with-intel-backing\/4236798\" target=\"_blank\" rel=\"nofollow noopener\">announced<\/a> in February, churning out 763 tokens a second in MiniMax M2.7 at short context lengths (10,000 input tokens) \u2014 several times faster than competing inference providers running on GPUs alone.<\/p>\n<p>Meanwhile, for longer context lengths, the company says that its platform is able to sustain more than 450 tokens a second.\u00a0<\/p>\n<p>                <img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/5268748.webp\" width=\"480\" height=\"274\" alt=\"Benchmark chart comparing SambaNova SN50 accelerators with GPU-only inference platforms.\" loading=\"lazy\" style=\"\"\/><\/p>\n<p>\n            At short context lengths (10,000 tokens) Artificial Analysis found SambaNova&#8217;s heterogeneous compute platform, which combined four H200 GPUs with 16 of its SN50 RDUs was able to achieve decode speed of 763 tok\/s in MiniMax M2.7.<br \/>\n            Artificial Analysis\n        <\/p>\n<p>This feat was accomplished by combining Nvidia GPUs with SambaNova Reconfigurable Dataflow Units (RDUs) to form a heterogeneous inference platform.<\/p>\n<p>Specifically, the computationally intensive prefill phase of the inference pipeline, during which prompts are processed and key value caches are generated, was handled by four Nvidia H200 GPUs.<\/p>\n<p>Meanwhile, memory-bandwidth-bound decode operations, where output tokens are generated, were done on a single SambaNova rack containing 16 SN50 accelerators.<\/p>\n<p>Disaggregating prefill from decode has become a key lever for reducing token costs for long-running AI agents, like code assistants. Nvidia initially <a href=\"https:\/\/www.theregister.com\/special-features\/2025\/03\/23\/a-closer-look-at-dynamo-nvidias-operating-system-for-ai\/1098518\" target=\"_blank\" rel=\"nofollow noopener\">demonstrated<\/a> this with its NVL72 rack systems, by varying the ratio of GPUs used for prefill versus decode. The company further disaggregated this with its <a href=\"https:\/\/www.theregister.com\/special-features\/2026\/03\/19\/a-closer-look-at-nvidias-groq-powered-lpx-rack-systems\/5223364\" target=\"_blank\" rel=\"nofollow noopener\">Groq-based LPX racks<\/a> revealed at GTC this spring. Since then, just about everyone from AMD to AWS and Cerebras has announced some kind of disaggregated or heterogeneous inference platform using one or more accelerators.<\/p>\n<p>With SambaNova&#8217;s latest performance figures, the startup hopes to demonstrate how customers can breathe new life into their aging GPU fleets by using its systems as decode accelerators. And because its systems are air-cooled, they can be deployed in existing datacenters \u2014 something that can&#8217;t be said of Nvidia&#8217;s latest generation of <a href=\"https:\/\/www.theregister.com\/on-prem\/2026\/01\/05\/nvidia-unpacks-vera-rubin-rack-system-at-ces\/2484391\" target=\"_blank\" rel=\"nofollow noopener\">Rubin GPUs<\/a>, which absolutely need liquid cooling.<\/p>\n<p>SambaNova plans to show off even more powerful inference configs, with 128 and eventually 256 accelerators to demonstrate its ability to maintain high token generation rates at high throughput. As we\u2019ve previously <a href=\"https:\/\/www.theregister.com\/on-prem\/2026\/03\/07\/unpacking-the-deceptively-simple-science-of-tokenomics\/5228181\" target=\"_blank\" rel=\"nofollow noopener\">explored<\/a>, this is something that GPUs alone have historically struggled with and one of the key drivers behind Nvidia\u2019s Groq <a href=\"https:\/\/www.theregister.com\/on-prem\/2025\/12\/31\/why-did-nvidia-really-drop-20b-on-groq\/2018042\" target=\"_blank\" rel=\"nofollow noopener\">acquihire<\/a> late last year.<\/p>\n<p>The results come just a month after SambaNova and Intel announced Vector Core Compute would be among the first to deploy the combined GPU + RDU offering with TogetherAI as their first large-scale customer.<\/p>\n<p>Ramping production of any chip isn\u2019t a cheap prospect, but for its fifth-gen part, capital shouldn\u2019t be an issue. On Wednesday, SambaNova <a href=\"https:\/\/sambanova.ai\/press\/sambanova-completes-first-close-of-1b-financing-at-11b-valuation?hs_amp=true\" rel=\"nofollow noopener\" target=\"_blank\">completed<\/a> the first close of a $1 billion Series F funding round led by General Atlantic, giving the AI chip startup an $11 billion valuation. \u00ae<\/p>\n","protected":false},"excerpt":{"rendered":"ai and ml Third-party testing shows heterogeneous compute platform combining H200s and SN50 RDUs churning out 763 tok\/s&hellip;\n","protected":false},"author":2,"featured_media":532246,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-532245","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/532245","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=532245"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/532245\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/532246"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=532245"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=532245"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=532245"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}