{"id":370153,"date":"2026-04-01T22:50:14","date_gmt":"2026-04-01T22:50:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/370153\/"},"modified":"2026-04-01T22:50:14","modified_gmt":"2026-04-01T22:50:14","slug":"nvidia-extreme-co-design-delivers-new-mlperf-inference-records","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/370153\/","title":{"rendered":"NVIDIA Extreme Co-Design Delivers New MLPerf Inference Records"},"content":{"rendered":"<p>Co-designed hardware, software, and models are key to delivering the highest AI factory throughput and lowest token cost. Measuring this goes far beyond peak chip specifications. Rigorous AI inference performance benchmarks are critical to understanding real-world token output, which drives AI factory revenue.<\/p>\n<p>MLPerf Inference v6.0 is the latest in a series of industry benchmarks that measure performance across a wide range of model architectures and use cases. In this latest round, systems powered by NVIDIA Blackwell Ultra GPUs delivered the highest throughput across the widest range of models and scenarios. This brings the cumulative NVIDIA MLPerf training and inference wins since 2018 to 291, which is 9x of all other submitters combined.<\/p>\n<p>This round, the NVIDIA partner ecosystem participated broadly, with 14 partners\u2014the largest number of partners submitting on any platform. ASUS, Cisco, <a href=\"https:\/\/www.coreweave.com\/blog\/coreweaves-innovation-velocity-drives-mlperf-6-0-leadership\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">CoreWeave<\/a>, Dell Technologies, GigaComputing, Google Cloud, <a href=\"https:\/\/community.hpe.com\/t5\/ai-unlocked\/hpe-secures-18-top-rankings-in-mlperf-inference-v6-0-ai\/ba-p\/7264016\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">HPE<\/a>, <a href=\"https:\/\/lenovopress.lenovo.com\/lp2413-lenovo-and-nvidia-advance-ai-training-efficiency-with-mlperf-60\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Lenovo<\/a>, <a href=\"https:\/\/nebius.com\/blog\/posts\/mlperf-inference-v6-0-results\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Nebius<\/a>, Netweb Technology, Quanta Cloud Technology (QCT), Red Hat, <a href=\"https:\/\/learn-more.supermicro.com\/data-center-stories\/supermicro-leads-whisper-benchmark-in-mlperf-v6-nvidia-b300-gpus\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Supermicro<\/a>, and <a href=\"https:\/\/lambda.ai\/blog\/lambdas-mlperf-inference-v6.0-hardware-leap-software-maturity-research-breakthrough?utm_source=linkedin&amp;utm_medium=organic-social&amp;utm_campaign=2026-04-mlperf-inference-v6&amp;utm_content=post-1\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Lambda<\/a> have delivered excellent performance on the NVIDIA platform.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"1125\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/04\/MLPerf-9x.webp.webp\" alt=\"Line chart showing NVIDIA GPU MLPerf training and inference wins growing from 2018 to 2026, accumulating 9x more wins than all others combined, alongside logos of NVIDIA partners delivering outstanding performance including ASUS, Cisco, CoreWeave, Dell, Google Cloud, HPE, Lenovo, Lambda, Nebius, Netweb, QCT, Red Hat, and Supermicro.\" class=\"lazyload wp-image-115044\"  data-\/><\/p>\n<p>\t\tFigure 1. NVIDIA Delivers 9x More Cumulative MLPerf Training and Inference wins<\/p>\n<p>This post takes a closer look at the latest benchmark updates, the industry-leading performance achieved on the NVIDIA platform, and the full-stack engineering that makes it possible.\u00a0<\/p>\n<p>New benchmarks, new performance records<a href=\"#new_benchmarks_new_performance_records\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>The MLPerf Inference benchmark suite is routinely updated to ensure that it reflects models, modalities, use cases, and deployment scenarios that matter to the community. Only the NVIDIA platform submitted results on all newly added models and scenarios this round, and delivered the highest performance across all of them.<\/p>\n<p>This round of MLPerf Inference added several new tests, including:<\/p>\n<p>DeepSeek-R1 Interactive: Following the addition of DeepSeek-R1 reasoning LLM based on a sparse mixture-of-experts (MoE) architecture in MLPerf Inference v5.1, MLCommons added a new Interactive scenario with 5x faster minimum token rate and 1.3x shorter time to first token compared to the server scenario, representing higher-interactivity deployments<\/p>\n<p><a href=\"https:\/\/mlcommons.org\/2026\/02\/vlm-inference-shopify\/\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Qwen3-VL-235B-A22B<\/a>: Vision-language model with a total of 235B parameters. This represents the first multi-modal model in the MLPerf Inference suite. Two scenarios are tested: Offline and Server.\u00a0<\/p>\n<p><a href=\"https:\/\/mlcommons.org\/2026\/03\/mlperf-inference-gpt-oss\/\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">GPT-OSS-120B<\/a>: 120B-parameter MoE reasoning LLM, developed by OpenAI. This benchmark includes three scenarios: Offline, Server, and Interactive<\/p>\n<p><a href=\"https:\/\/mlcommons.org\/2026\/03\/texttovideo-inference\/\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">WAN-2.2-T2V-A14B<\/a>: 4B-parameter text-to-video generative AI model. Two scenarios tested: single-stream, which measures the latency to process a single video generation request, and offline, which measures the number of samples processed per second in a batch-processing scenario.\u00a0<\/p>\n<p><a href=\"https:\/\/mlcommons.org\/2026\/02\/dlrmv3-inference-meta\/\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">DLRMv3 <\/a>\u2013 A generative recommendation benchmark that replaces the DLRM-DCNv2 test. It uses a transformer-based architecture that increases model size and compute intensity compared to the prior benchmark. It tests offline and server scenarios.\u00a0<\/p>\n<p>BenchmarkDeepSeek-R1GPT-OSS-120BQwen3-VLWan 2.2DLRMv3Offline2,494,310 tokens\/sec*1,046,150 tokens\/sec79 samples\/sec0.059 samples\/sec104,637 samples\/secServer1,555,110 tokens\/sec*1,096,770 tokens\/sec68 queries\/sec21 secs**(Single Stream)99,997 queries\/secInteractive250,634 tokens\/sec677,199 tokens\/sec*********Table 1. NVIDIA platform throughput on newly added workloads and scenarios in MLPerf Inference v6.0\u00a0<\/p>\n<p class=\"has-small-font-size\">* Not a new scenario in MLPerf Inference v6.0<br \/>** Wan 2.2 features a single stream scenario, which measures end-to-end request latency, instead of a server scenario. Lower is better.<br \/>*** Not tested in MLPerf Inference v6.0<\/p>\n<p class=\"has-text-color has-link-color has-small-font-size wp-elements-6b26afad7efab72e15d1e41abea634f4\" style=\"color:#605e5e\">MLPerf Inference v6.0, Closed Division. Results retrieved from www.mlcommons.org on April 1, 2026. NVIDIA platform results from the following entries: 6.0-0039, 6.0-0073, 6.0-0075, 6.0-0076, 6.0-0078, 6.0-0081, 6.0-0094. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" width=\"1999\" height=\"1125\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on-async--click=\"actions.showLightbox\" data-wp-on-async--load=\"callbacks.setButtonStyles\" data-wp-on-async-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/04\/MLPerf-2.7x.webp.webp\" alt=\"Bar chart showing 2.7x DeepSeek-R1 inference improvement on NVIDIA GB300 NVL72 between August 2025 and February 2026, alongside a record 2.5 million tokens per second achieved on 288 NVIDIA Blackwell Ultra GPUs.\" class=\"lazyload wp-image-115046\"  data-\/><\/p>\n<p>\t\tFigure 2. NVIDIA achieves 2.7x performance gain and 2.5M token\/s on DeepSeek-R1<\/p>\n<p>NVIDIA TensorRT-LLM software updates unlock up to 2.7X performance gains on the same Blackwell Ultra GPUs<a href=\"#nvidia_tensorrt-llm_software_updates_unlock_up_to_27x_performance_gains_on_the_same_blackwell_ultra_gpus\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>NVIDIA continually optimizes the performance of its software stack to increase delivered token throughput from existing platforms. This delivers reductions in token production cost and enables AI factory operators to serve more users to generate more revenue with a given infrastructure footprint.<\/p>\n<p>The additional performance also provides headroom to run future AI models and serve existing models in demanding scenarios, such as higher token rates and longer contexts. This continual improvement makes it possible for NVIDIA GPUs introduced years ago to remain productive, at high utilization rates, in the cloud.\u00a0<\/p>\n<p>This round, NVIDIA GB300 NVL72\u2014launched last year\u2014delivered up to 2.7x higher token throughput compared to its debut submissions just six months ago on the server scenario of the DeepSeek-R1 benchmark1. This means 2.7x more tokens from the same GB300 NVL72-based infrastructure and power footprint, reducing the cost to manufacture each token by more than 60%. This speedup, achieved by NVIDIA partner <a href=\"https:\/\/nebius.com\/blog\/posts\/mlperf-inference-v6-0-results\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">Nebius<\/a>, showcases a core advantage of the NVIDIA platform: an open, expansive ecosystem where customers and partners can uniquely optimize and innovate on top of our software stack.\u00a0<\/p>\n<p class=\"has-text-color has-link-color has-small-font-size wp-elements-a68a1a62ec5e6d784c62d76a30427d14\" style=\"color:#605e5e\">1MLPerf Inference v5.1 and v6.0, Closed Division. Results retrieved from <a href=\"http:\/\/www.mlcommons.org\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">www.mlcommons.org<\/a> on April 1, 2026. NVIDIA platform results from the following entries: 5.1-0072,\u00a0 6.0-0081. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See <a href=\"http:\/\/www.mlcommons.org\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">www.mlcommons.org<\/a> for more information.<\/p>\n<p>Powering the DeepSeek R1 performance improvements in the server and offline scenarios were several software enhancements, including:<\/p>\n<p>Faster kernels\u2014this included a combination of higher-performance kernels and the use of fewer kernels because of kernel fusions.<\/p>\n<p>Optimized Attention Data Parallel\u2014Better balancing of context requests between different ranks, enabling significant speedups in end-to-end performance.<\/p>\n<p>The latest features of the open source NVIDIA TensorRT-LLM inference serving software and the <a href=\"https:\/\/developer.nvidia.com\/blog\/nvidia-dynamo-1-production-ready\/\" data-wpel-link=\"internal\" target=\"_self\" rel=\"follow nofollow noopener\">NVIDIA Dynamo<\/a> open source distributed inference serving framework were used to support the newly added and more challenging DeepSeek-R1 Interactive scenario. This includes:<\/p>\n<p>Disaggregated serving: This capability in Dynamo separates and individually optimizes the configurations of each inference phase (prefill and decode), respectively, enabling optimal overall throughput.\u00a0<\/p>\n<p>Wide Expert Parallel (WideEP): For higher-interactivity scenarios, execution time for MoE models is bound by expert weight load time. By splitting, or sharding, the experts across multiple GPUs across NVL72 nodes, this bottleneck is reduced, improving end-to-end performance.<\/p>\n<p>Multi-Token Prediction (MTP): At higher interactivity levels, batch sizes are smaller, and performance is dominated by how quickly weights can load into memory, leaving compute performance underutilized. By applying compute otherwise that goes unutilized to predict and verify additional tokens in parallel (up to three in this implementation), throughput at high interactivity is increased.\u00a0<\/p>\n<p>KV-aware routing: This capability of Dynamo routes inference requests by evaluating their compute costs across different workers.<\/p>\n<p>NVIDIA was the first and only platform to submit DeepSeek-R1 results on MLPerf Inference when the benchmark debuted last year. This round, NVIDIA not only increased performance on returning scenarios for DeepSeek-R1 but \u200cwas once again the only platform to submit on the newly added interactive scenario.<\/p>\n<p>And even on Llama 3.1 405B\u2014a very large, dense LLM launched almost two years ago\u2014 GB300 NVL72 performance increased by 1.5x in the server scenario.<\/p>\n<p>BenchmarkGB300 NVL72\u00a0v5.1GB300 NVL72v6.0SpeedupDeepSeek-R1(Server)2,907 tokens\/sec\/gpu8,064 tokens\/sec\/gpu2.77xDeepSeek-R1(Offline)5,842 tokens\/sec\/gpu9,821 tokens\/sec\/gpu1.68xLlama 3.1 405B(Server)170 tokens\/sec\/gpu259 tokens\/sec\/gpu1.52xLlama 3.1 405B(Offline)224 tokens\/sec\/gpu271 tokens\/sec\/gpu1.21xTable 2. Performance improvements, normalized on a per-GPU basis, on DeepSeek-R1 and Llama 3.1 405B server and offline scenarios in v6.0 compared to v5.1<\/p>\n<p class=\"has-text-color has-link-color has-small-font-size wp-elements-4e322fc567939449d5f8e068d2d4b507\" style=\"color:#605e5e\">MLPerf Inference v5.1 and v6.0, Closed Division. Results retrieved from www.mlcommons.org on April 1, 2026. NVIDIA platform results from the following entries: 5.1-0072, 6.0-0017, 6.0-0078, 6.0-0082. Per chip performance is derived by dividing total throughput by the number of reported chips. Per-chip performance is not a primary metric of MLPerf Inference v5.1 or v6.0. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.<\/p>\n<p>Additionally, NVIDIA submissions on the newly added multimodal, video generation, and recommendation benchmarks were powered by open source software frameworks optimized for the NVIDIA platform. The Qwen3-VL vision-language submission used the vLLM open source framework, showing how the community is rapidly building advanced multimodal optimizations to accelerate image-heavy inference workloads on the latest GPUs like NVIDIA Blackwell Ultra. The WAN-2.2 text-to-video submission used the <a href=\"https:\/\/github.com\/NVIDIA\/TensorRT-LLM\/tree\/main\/examples\/visual_gen\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">TensorRT-LLM VisualGen<\/a>, which accelerates diffusion-based video generation pipelines on NVIDIA GPUs.\u00a0<\/p>\n<p>For DLRMv3, the submission was built on two open-source projects: the NVIDIA <a href=\"https:\/\/github.com\/NVIDIA\/recsys-examples\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">recsys-example<\/a> for high-performance transformer-based recommendation inference, and <a href=\"https:\/\/github.com\/NVIDIA\/nv-embedding-cache\" data-wpel-link=\"external\" target=\"_blank\" rel=\"follow nofollow noopener\">NV Embedding Cache<\/a> for GPU-accelerated embedding table lookups. Both were critical to achieving record throughput on this more demanding generative recommendation benchmark.<\/p>\n<p>Through extensive and ongoing engineering, NVIDIA is continually increasing performance on existing hardware on existing models, as evidenced by these results. At the same time, NVIDIA collaborates closely with model builders and open source inference frameworks to ensure that the latest models run on the NVIDIA platform on the day of launch.\u00a0<\/p>\n<p>Scale-out inference with NVIDIA Quantum-X800 InfiniBand platform enables millions of tokens per second<a href=\"#scale-out_inference_with_nvidia_quantum-x800_infiniband_platform_enables_millions_of_tokens_per_second\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>NVIDIA also set new throughput records at scale on the DeepSeek-R1 model in the offline and server scenarios by submitting results using four GB300 NVL72 systems interconnected with <a href=\"https:\/\/www.nvidia.com\/en-us\/networking\/products\/infiniband\/quantum-x800\/\" data-wpel-link=\"internal\" target=\"_self\" rel=\"follow nofollow noopener\">NVIDIA Quantum-X800 InfiniBand<\/a> scale-out networking.\u00a0<\/p>\n<p>DeepSeek-R1 | 4x GB300 NVL72Tokens\/SecondOffline2,494,310Server1,555,110Table 3. DeepSeek-R1 throughput on four GB300 NVL72 systems scaled up with NVLink and scaled out with NVIDIA Quantum-X800 InfiniBand<\/p>\n<p class=\"has-text-color has-link-color has-small-font-size wp-elements-684de1540e2f218276b996e3adcc9224\" style=\"color:#605e5e\">MLPerf Inference v6.0, Closed Division. Results retrieved from www.mlcommons.org on April 1, 2026. NVIDIA platform results from the following entries: 6.0-0076. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.<\/p>\n<p>With 288 Blackwell Ultra GPUs\u2014the largest scale ever submitted to any benchmark in MLPerf Inference\u2014submissions set new system-level throughput records, enabling millions of tokens processed per second.\u00a0<\/p>\n<p>Looking ahead to MLPerf Endpoints\u00a0<a href=\"#looking_ahead_to_mlperf_endpoints\u00a0\" class=\"heading-anchor-link\"><\/a><\/p>\n<p>Delivered inference throughput takes extreme co-design across many chips, system architecture, data center design, and software. The latest MLPerf Inference v6.0 results show that NVIDIA yields unmatched inference throughput across the broadest range of workloads, from massive LLMs to advanced vision language models, to generative recommender systems and more, on industry-standard benchmarks.\u00a0<\/p>\n<p>AI inference workloads also continue to evolve rapidly, as model sizes grow and context lengths rise. As agentic AI becomes more prevalent, premium use cases that require ultra-fast token rates are emerging.\u00a0<\/p>\n<p>NVIDIA has been working, as part of the MLCommons consortium, to lead the definition of the MLPerf Endpoints benchmark. MLPerf Endpoints will give the community a rigorous, auditable picture of how deployed services perform under real API traffic\u2014capturing key performance metrics that chip-level benchmarks alone cannot reveal\u2014while providing the rigor and result integrity that defines MLPerf benchmarks.\u00a0<\/p>\n<p>To explore the latest performance on the NVIDIA platform across training, inference, and high-performance computing, please see our <a href=\"https:\/\/developer.nvidia.com\/deep-learning-performance-training-inference?sortBy=developer_learning_library%2Fsort%2Ffeatured_in.deep_learning_performance%3Adesc%2Ctitle%3Aasc\" data-wpel-link=\"internal\" target=\"_self\" rel=\"follow nofollow noopener\">deep learning product performance page<\/a>.\u00a0<\/p>\n<p>Acknowledgements<\/p>\n<p>NVIDIA MLPerf Inference v6.0 results reflect the work of many talented engineers across the company. We\u2019d like to acknowledge the contributions of the following individuals (last name sorted):<\/p>\n<p>Tomar Bar-on, Nitin Sai Bommi, Viraat Chandra, Alice Cheng, Jerry Chen, Xiaoming Chen, Jesus Corbal San Adrian, Ashutosh Dhar, Kefeng Duan, Wookje Han, Kyle Huang, Kris Hung, Rashid Kaleem, Khubaib Khubaib, Zihao Kong, Tin-Yin Lai, Tao Li, Forrest Lin, Wanqian Li, Alex Liu, Jintao Peng, Yuxian Qiu, Junyi Qiu, Xiaowei Shi, Olivia Stoner, Jacob Subag, Tong Tong, Harshil Vagadia, Shobhit Verma, June Yang, Tailing Yuan, Ben Zhang\u2026 and many others across NVIDIA whose efforts made these results possible.<\/p>\n","protected":false},"excerpt":{"rendered":"Co-designed hardware, software, and models are key to delivering the highest AI factory throughput and lowest token cost.&hellip;\n","protected":false},"author":2,"featured_media":370154,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-370153","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/370153","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=370153"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/370153\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/370154"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=370153"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=370153"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=370153"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}