{"id":322367,"date":"2026-02-28T20:19:19","date_gmt":"2026-02-28T20:19:19","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/322367\/"},"modified":"2026-02-28T20:19:19","modified_gmt":"2026-02-28T20:19:19","slug":"even-frontier-llms-from-gpt-5-onward-lose-up-to-33-accuracy-when-you-chat-too-long","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/322367\/","title":{"rendered":"Even frontier LLMs from GPT-5 onward lose up to 33% accuracy when you chat too long"},"content":{"rendered":"<p>The latest generation of large language models\u2014from GPT-5 onward\u2014still struggles when tasks are spread across multiple conversation turns. Researcher Philippe Laban and his team tested current models on six tasks covering code, databases, actions, data-to-text, math, and summarization. Performance drops significantly when information is split across several messages (sharded) instead of a single prompt (concat).<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/02\/lost_in_conversation_2026.jpg\"><img data-lazyloaded=\"1\" fetchpriority=\"high\" decoding=\"async\" class=\"wp-image-32652 size-full\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/02\/lost_in_conversation_2026.jpg\" alt=\"\" width=\"959\" height=\"835\"\/><\/a><a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/x.com\/PhilippeLaban\/status\/2026329136864645390\">Laban et al.<\/a><\/p>\n<p>Newer models did slightly better\u2014performance degradation shrank from 39 to 33 percent\u2014but the issue is far from solved. The biggest gains showed up in Python tasks, where some models only lost 10 to 20 percent. Laban suspects real-world losses could be even worse, since the tests used simple user simulations. Users who change their mind mid-conversation would likely cause steeper drops.<\/p>\n<p>Technical tweaks like lowering temperature values don&#8217;t fix the problem, <a href=\"https:\/\/the-decoder.com\/ai-chatbots-become-dramatically-less-reliable-in-longer-conversationsnew-study-finds\/\" rel=\"nofollow noopener\" target=\"_blank\">the original study found<\/a>. The researchers recommend starting a fresh conversation when things go sideways, ideally by having the model summarize all requests first and using that summary as the starting point for a new chat.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tAs a THE DECODER subscriber, you get ad-free reading, our weekly AI newsletter, the exclusive &#8220;AI Radar&#8221; Frontier Report 6\u00d7 per year, access to comments, and our complete archive.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"The latest generation of large language models\u2014from GPT-5 onward\u2014still struggles when tasks are spread across multiple conversation turns.&hellip;\n","protected":false},"author":2,"featured_media":322368,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[149780,61,60,5342,80],"class_list":["post-322367","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-context-engineering","tag-ie","tag-ireland","tag-prompt-engineering","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/322367","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=322367"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/322367\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/322368"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=322367"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=322367"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=322367"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}