{"id":661688,"date":"2026-06-27T20:44:18","date_gmt":"2026-06-27T20:44:18","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/661688\/"},"modified":"2026-06-27T20:44:18","modified_gmt":"2026-06-27T20:44:18","slug":"accelerating-gemini-nano-models-on-pixel-with-frozen-multi-token-prediction","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/661688\/","title":{"rendered":"Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction"},"content":{"rendered":"<p data-block-key=\"2mxvd\">Having powerful Large Language Models (LLMs) right in your pocket is now a reality with on-device models like <a href=\"https:\/\/developer.android.com\/ai\/gemini-nano\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Gemini Nano<\/a> and <a href=\"https:\/\/deepmind.google\/models\/gemma\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Gemma<\/a>. This technology enables everyday features on your phone \u2014 such as instantly summarizing a flurry of notifications or proofreading an important text message \u2014 all without sending your private data off device. But to make these features useful for everyday users, they need to happen very efficiently.<\/p>\n<p data-block-key=\"dq4tq\">Delivering this kind of speed on a mobile device is a significant challenge. Unlike vast server environments, mobile phones operate under a strict energy budget and hard memory (RAM) limits. Furthermore, standard language models generate text &#8220;autoregressively&#8221; \u2014 meaning they process and output just one word (or token) at a time. This step-by-step process creates a bottleneck, underutilizing the phone&#8217;s processing power while straining its memory bandwidth, which can ultimately slow down the user experience and drain the battery.<\/p>\n<p data-block-key=\"baqq1\">To overcome this bottleneck, we are announcing a new architecture that retrofits Multi-Token Prediction (MTP) onto existing, &#8220;frozen&#8221; Gemini Nano v3 models. Building on prior approaches like the<a href=\"https:\/\/arxiv.org\/pdf\/2401.15077\" target=\"_blank\" rel=\"noopener noreferrer nofollow\"> EAGLE framework<\/a> and <a href=\"https:\/\/research.google\/blog\/accelerating-text-generation-with-confident-adaptive-language-modeling-calm\/\" rel=\"nofollow noopener\" target=\"_blank\">Confident Adaptive Language Modeling<\/a> (CALM), we designed new architectural components to maximize these efficiency gains specifically for mobile environments. Our recent announcements highlighted accelerating <a href=\"https:\/\/blog.google\/innovation-and-ai\/technology\/developers-tools\/multi-token-prediction-gemma-4\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Gemma 4 with MTP<\/a> and making it available to developers.<\/p>\n<p data-block-key=\"c6tua\">Today&#8217;s article tackles the unique, extreme constraints of edge computing. Recently rolled out to the Pixel 9 and 10 series, this approach acts as an out-of-the-box speedup. For users, this means that features like AI Notification Summaries and Proofread generate text significantly faster and with less energy consumption. For developers, it eliminates a major friction point: delivering high-speed on-device AI without the need to fine-tune separate, memory-heavy drafting models for every new task.<\/p>\n","protected":false},"excerpt":{"rendered":"Having powerful Large Language Models (LLMs) right in your pocket is now a reality with on-device models like&hellip;\n","protected":false},"author":2,"featured_media":648401,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-661688","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/661688","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=661688"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/661688\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/648401"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=661688"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=661688"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=661688"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}