{"id":495187,"date":"2026-06-12T00:46:22","date_gmt":"2026-06-12T00:46:22","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/495187\/"},"modified":"2026-06-12T00:46:22","modified_gmt":"2026-06-12T00:46:22","slug":"googles-latest-diffusiongemma-open-ai-model-comes-with-a-4x-speed-boost","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/495187\/","title":{"rendered":"Google&#8217;s latest DiffusionGemma open AI model comes with a 4x speed boost"},"content":{"rendered":"<p>Another day, another AI model from Google. This time, Google DeepMind has released a new member of the <a href=\"https:\/\/arstechnica.com\/ai\/2026\/04\/google-announces-gemma-4-open-ai-models-switches-to-apache-2-0-license\/\" rel=\"nofollow noopener\" target=\"_blank\">Gemma 4 open model family<\/a>, but it\u2019s fundamentally different from the rest of the lineup. DiffusionGemma doesn\u2019t generate outputs linearly like most AI models. Instead, it can produce an entire block of text in parallel. <a href=\"https:\/\/blog.google\/innovation-and-ai\/technology\/developers-tools\/diffusion-gemma-faster-text-generation\/\" rel=\"nofollow noopener\" target=\"_blank\">Google says<\/a> this makes it faster and more efficient when running on local hardware like an Nvidia DGX or a humble gaming GPU.<\/p>\n<p>Most AI models are designed to be autoregressive\u2014they generate text left to right one token at a time. DiffusionGemma has more in common with image generation models, which start with static and then denoise it to create the desired content. This model takes a field of placeholder tokens running over the canvas multiple times to generate likely tokens and using those to improve estimation of others. At the end of the process, the model finalizes its token outputs in one large block\u2014the \u201cdenoised\u201d text canvas.<\/p>\n<p>DiffusionGemma is fairly large in the realm of Google\u2019s open models. It\u2019s a Mixture of Experts (MoE) model with a total of 26 billion parameters, but only 3.8 billion are activated during inference. That means it should fit in the 18GB RAM allotment of a high-end GPU. In testing with an RTX 5090, DiffusionGemma spits out around 700 tokens per second. With a single Nvidia H100 AI accelerator, DiffusionGemma can produce 1,000+ tokens per second. That\u2019s about four times the output of the similarly sized autoregressive Gemma models.<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/06\/sudoku_before_after11.gif\"><\/p>\n<p>            <a class=\"cursor-zoom-in\" data-pswp-width=\"1586\" data-pswp-height=\"948\" data-cropped=\"false\" href=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/06\/sudoku_before_after11.gif\" target=\"_blank\" data-pswp-><br \/>\n              <img width=\"1586\" height=\"948\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/06\/sudoku_before_after11.gif\" class=\"fullwidth full\" alt=\"\" decoding=\"async\" loading=\"lazy\"\/><br \/>\n            <\/a><\/p>\n<p><\/a><\/p>\n<p>This approach to text generation shifts the bottleneck from memory bandwidth to compute, generating up to 256 tokens in parallel. Google says this offers a measurable boost in non-linear tasks like in-line editing, molecular sequencing, and mathematical graphing. The animation above shows how DiffusionGemma was tuned to solve Sudoku puzzles, which is a notoriously challenging task for standard autoregressive AI models because each token depends on future tokens. DiffusionGemma\u2019s ability to continuously self-correct large sets of tokens makes that easier.<\/p>\n","protected":false},"excerpt":{"rendered":"Another day, another AI model from Google. This time, Google DeepMind has released a new member of the&hellip;\n","protected":false},"author":2,"featured_media":495188,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[220,218,219,61,60,80],"class_list":["post-495187","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-ie","tag-ireland","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/495187","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=495187"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/495187\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/495188"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=495187"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=495187"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=495187"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}