{"id":418421,"date":"2026-04-30T13:43:11","date_gmt":"2026-04-30T13:43:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/418421\/"},"modified":"2026-04-30T13:43:11","modified_gmt":"2026-04-30T13:43:11","slug":"google-ai-breakthrough-means-chatbots-use-six-times-less-memory-during-conversations-without-compromising-performance","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/418421\/","title":{"rendered":"Google AI breakthrough means chatbots use six times less memory during conversations without compromising performance"},"content":{"rendered":"<p id=\"elk-d4585d0e-ca4a-4beb-a11b-edf6fc7d7e33\">Google engineers have developed a method to compress <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\" data-url=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\" rel=\"nofollow noopener\" target=\"_blank\">artificial intelligence<\/a> (AI) data so that it requires up to six times less working memory to function.<\/p>\n<p>With the new system, called TurboQuant, AI algorithms could retain the same amount of information and perform equally powerful computations, but with significantly less memory hardware, the company says.<\/p>\n<p><a id=\"elk-seasonal\"\/><\/p>\n<p id=\"elk-d4585d0e-ca4a-4beb-a11b-edf6fc7d7e33-2\" class=\"paywall\" aria-hidden=\"true\">Current AI algorithms need a lot of working memory, also known as the key value (KV) cache, to work properly. This is where immediate computational results and other bits of info are stored temporarily during active processing.<\/p>\n<p>            You may like<\/p>\n<p id=\"elk-289ba9ce-0600-49e1-913c-5e86f53dbcb9\">For example, if you ask ChatGPT what the weather will be like tomorrow in your area, it may store words like &#8220;weather&#8221; and &#8220;tomorrow,&#8221; along with your location and partial guesses, like &#8220;It might be rainy,&#8221; in the KV cache while it generates its response. The larger an AI model&#8217;s KV cache is, the more information it can keep track of at once and the more powerful it is.<\/p>\n<p>A single sentence uses only a few dozen <a data-analytics-id=\"inline-link\" href=\"https:\/\/help.openai.com\/en\/articles\/4936856-what-are-tokens-and-how-to-count-them\" target=\"_blank\" data-url=\"https:\/\/help.openai.com\/en\/articles\/4936856-what-are-tokens-and-how-to-count-them\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">tokens<\/a> \u2014 the building blocks of AI prompts and output text \u2014 but storing hundreds of thousands of tokens in the KV cache for more sophisticated work <a data-analytics-id=\"inline-link\" href=\"https:\/\/developer.nvidia.com\/blog\/accelerate-large-scale-llm-inference-and-kv-cache-offload-with-cpu-gpu-memory-sharing\/\" target=\"_blank\" data-url=\"https:\/\/developer.nvidia.com\/blog\/accelerate-large-scale-llm-inference-and-kv-cache-offload-with-cpu-gpu-memory-sharing\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">can require tens of gigabytes of memory<\/a>. These memory requirements scale linearly depending on the number of users, and ChatGPT is known to receive <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-do-ai-chatbots-use-so-much-energy\" data-url=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-do-ai-chatbots-use-so-much-energy\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-do-ai-chatbots-use-so-much-energy\" rel=\"nofollow noopener\" target=\"_blank\">billions of requests<\/a> every day.<\/p>\n<p>The compression algorithm will decrease the amount of working memory an AI model needs to perform the same computations. It does so via a process called quantization, which results in values represented by fewer bits.<\/p>\n<p>Although Google has been using quantization on its neural networks for many years, it has typically been applied statically \u2014 that is, the compression is done once and doesn&#8217;t change as the model runs. The difference with TurboQuant is that it reduces the KV cache&#8217;s memory in real time \u202a\u2014\u202c a tricky feat given that it must keep the quantized data in the cache accurate and up-to-date while the model generates outputs.<\/p>\n<p class=\"newsletter-form__strapline\">Get the world\u2019s most fascinating discoveries delivered straight to your inbox.<\/p>\n<p>In a <a data-analytics-id=\"inline-link\" href=\"https:\/\/research.google\/blog\/turboquant-redefining-ai-efficiency-with-extreme-compression\/\" target=\"_blank\" data-url=\"https:\/\/research.google\/blog\/turboquant-redefining-ai-efficiency-with-extreme-compression\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">statement<\/a>, Google representatives said TurboQuant &#8220;showed great promise for reducing key-value bottlenecks without sacrificing AI model performance&#8221; in tests in Meta&#8217;s Llama 3.1-8B, Google&#8217;s Gemma and Mistral AI models.<\/p>\n<p>&#8220;This has potentially profound implications for all compression-reliant use cases, including and especially in the domains of search and AI,&#8221; they added.<\/p>\n<p><a id=\"elk-01c75a75-3803-4e12-9b2d-3845901c8297\" class=\"paywall\" aria-hidden=\"true\"\/>Is this Google&#8217;s &#8220;DeepSeek moment&#8221;?<\/p>\n<p id=\"elk-d2a1a120-bea9-414f-b762-e0a262024a19\">Google says TurboQuant could reduce the KV cache&#8217;s size by a factor of at least six times, using two methods: <a data-analytics-id=\"inline-link\" href=\"https:\/\/arxiv.org\/abs\/2502.02617\" target=\"_blank\" data-url=\"https:\/\/arxiv.org\/abs\/2502.02617\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">PolarQuant<\/a> and <a data-analytics-id=\"inline-link\" href=\"https:\/\/dl.acm.org\/doi\/10.1609\/aaai.v39i24.34773\" target=\"_blank\" data-url=\"https:\/\/dl.acm.org\/doi\/10.1609\/aaai.v39i24.34773\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">Quantized Johnson-Lindenstrauss<\/a> (QJL).<\/p>\n<p>            What to read next<\/p>\n<p>To interpret these methods, it is important to understand that data in the AI&#8217;s working memory has been turned into vectors \u2014 groups of numbers that have a defined size (radius) and direction (angle). Vectors can be mathematically &#8220;rotated,&#8221; meaning they are reexpressed in a different, common coordinate system.<\/p>\n<p>PolarQuant quantization reexpresses AI data from Cartesian coordinates (along X, Y and Z axes) into polar coordinates (angles around a single point). The rotation aligns the angles of the vectors more consistently, thereby allowing them to be compressed into fewer bits with less additional scaling information. The vectors then go through the QJL optimization method, where they are adjusted very slightly to correct any computational errors stemming from the quantization.<\/p>\n<p>In a <a data-analytics-id=\"inline-link\" href=\"https:\/\/x.com\/eastdakota\/status\/2036827179150168182\" target=\"_blank\" data-url=\"https:\/\/x.com\/eastdakota\/status\/2036827179150168182\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow\">post on the social media platform X<\/a>, <a data-analytics-id=\"inline-link\" href=\"https:\/\/blog.cloudflare.com\/author\/matthew-prince\/\" data-url=\"https:\/\/blog.cloudflare.com\/author\/matthew-prince\/\" target=\"_blank\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">Matthew Prince<\/a>, CEO of web security company Cloudflare, called the compression breakthrough &#8220;<a data-analytics-id=\"inline-link\" href=\"https:\/\/x.com\/eastdakota\/status\/2036827179150168182\" target=\"_blank\" data-url=\"https:\/\/x.com\/eastdakota\/status\/2036827179150168182\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow\">Google&#8217;s DeepSeek<\/a>&#8221; \u202a\u2014\u202c a reference to the surprise release of the Chinese firm&#8217;s AI model that <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-is-deekspeek-such-a-game-changer-scientists-explain-how-the-ai-models-work-and-why-they-were-so-cheap-to-build\" data-url=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-is-deekspeek-such-a-game-changer-scientists-explain-how-the-ai-models-work-and-why-they-were-so-cheap-to-build\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" data-before-rewrite-localise=\"https:\/\/www.livescience.com\/technology\/artificial-intelligence\/why-is-deekspeek-such-a-game-changer-scientists-explain-how-the-ai-models-work-and-why-they-were-so-cheap-to-build\" rel=\"nofollow noopener\" target=\"_blank\">achieved comparable results<\/a> to leading chatbots at a fraction of the cost.<\/p>\n<p>Google&#8217;s March 24 unveiling of TurboQuant sent stocks in memory companies <a data-analytics-id=\"inline-link\" href=\"https:\/\/uk.investing.com\/news\/stock-market-news\/mu-wdc-sndk-fall-why-googles-turboquant-is-rattling-memory-stocks-4576725\" target=\"_blank\" data-url=\"https:\/\/uk.investing.com\/news\/stock-market-news\/mu-wdc-sndk-fall-why-googles-turboquant-is-rattling-memory-stocks-4576725\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">like SanDisk, Western Digital and Seagate<\/a> plummeting. But although the discovery could prove pivotal in improving AI efficiency, it is still at the lab stage and has yet to be widely rolled out in real-world models.<\/p>\n<p id=\"elk-03c98fd0-903f-46e4-87a1-a57c7565d1d3\">Moreover, it will compress only the working memory used during inference. This is when it is generating a response to a prompt. A model&#8217;s training typically requires <a data-analytics-id=\"inline-link\" href=\"https:\/\/training.continuumlabs.ai\/infrastructure\/data-and-memory\/calculating-gpu-memory-for-serving-llms\" target=\"_blank\" data-url=\"https:\/\/training.continuumlabs.ai\/infrastructure\/data-and-memory\/calculating-gpu-memory-for-serving-llms\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">up to four times more memory<\/a> than that, so the actual impact on memory will be relatively small.<\/p>\n<p>This is what Merrill Lynch banker Vivek Arya explained to concerned investors in a note, according to <a data-analytics-id=\"inline-link\" href=\"https:\/\/www.zdnet.com\/article\/what-googles-turboquant-can-and-cant-do-for-ais-spiraling-cost\/\" target=\"_blank\" data-url=\"https:\/\/www.zdnet.com\/article\/what-googles-turboquant-can-and-cant-do-for-ais-spiraling-cost\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">ZDNet<\/a>: &#8220;(The) 6x improvement in memory efficiency [will] likely [lead] to 6x increase in accuracy (model size) and\/or context length (KV cache allocation), rather than 6x decrease in memory.&#8221;<\/p>\n<p>Google officially unveiled TurboQuant at <a data-analytics-id=\"inline-link\" href=\"https:\/\/iclr.cc\/\" target=\"_blank\" data-url=\"https:\/\/iclr.cc\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">ICLR 2026<\/a>, which took place April 23-27 in Rio de Janeiro, and will formally present PolarQuant and QJL at <a data-analytics-id=\"inline-link\" href=\"https:\/\/virtual.aistats.org\/\" target=\"_blank\" data-url=\"https:\/\/virtual.aistats.org\/\" referrerpolicy=\"no-referrer-when-downgrade\" data-hl-processed=\"none\" data-mrf-recirculation=\"inline-link\" rel=\"nofollow noopener\">AISTATS 2026<\/a> in Tangier, Morocco, in early May.<\/p>\n","protected":false},"excerpt":{"rendered":"Google engineers have developed a method to compress artificial intelligence (AI) data so that it requires up to&hellip;\n","protected":false},"author":2,"featured_media":418422,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-418421","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/418421","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=418421"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/418421\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/418422"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=418421"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=418421"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=418421"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}