{"id":652809,"date":"2026-06-22T20:39:12","date_gmt":"2026-06-22T20:39:12","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/652809\/"},"modified":"2026-06-22T20:39:12","modified_gmt":"2026-06-22T20:39:12","slug":"glm-5-2-is-the-step-change-for-open-agents","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/652809\/","title":{"rendered":"GLM-5.2 is the step change for open agents"},"content":{"rendered":"<p>Housekeeping: Following my \u201c<a href=\"https:\/\/www.interconnects.ai\/p\/state-of-the-blog-mid-2026\" rel=\"nofollow noopener\" target=\"_blank\">State of the blog<\/a>\u201d post last week, noting a slight increase in paid features, it\u2019s a good time to remind folks that I offer <a href=\"https:\/\/www.interconnects.ai\/about#\u00a7group-paid-subscriptions\" rel=\"nofollow noopener\" target=\"_blank\">group subscriptions<\/a> with larger discounts proportional to the number of seats. <br \/>I also released a new paper today on open RL recipes for terminal agents, read more <a href=\"https:\/\/natolambert.substack.com\/p\/tmax-an-open-rl-recipe-for-terminal\" rel=\"nofollow noopener\" target=\"_blank\">here<\/a>.<\/p>\n<p>A bit over a week ago, when the AI world was still reeling from the shocking <a href=\"https:\/\/www.interconnects.ai\/p\/welcome-to-the-agi-era-of-ai-governance\" rel=\"nofollow noopener\" target=\"_blank\">export restriction, and effective banning<\/a>, of <a href=\"https:\/\/www.interconnects.ai\/p\/claude-fable-5-and-new-ai-safety\" rel=\"nofollow noopener\" target=\"_blank\">Claude Fable 5<\/a>, Z.ai released their latest model, GLM-5.2. This model was <a href=\"https:\/\/x.com\/Zai_org\/status\/2065704919299235870\" rel=\"nofollow\">rolled out<\/a> unusually on a Saturday, June 13th, to GLM Coding Plan members. This is an unusual release practice, normally when an AI model is released on a weekend it\u2019s for a weird reason (most famously, <a href=\"https:\/\/www.interconnects.ai\/p\/llama-4\" rel=\"nofollow noopener\" target=\"_blank\">Llama 4<\/a>). In this case, it seemed like Z.ai was excited to capitalize on the zeitgeist of \u201cAnthropic being anti open-science\u201d with their silent safeguards on AI researchers. For the past year or two, the Chinese open-weight labs have taken every opportunity they have for easy marketing wins like this.<\/p>\n<p data-attrs=\"{&quot;url&quot;:&quot;https:\/\/www.interconnects.ai\/p\/glm-52-is-the-step-change-for-open?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}\" data-component-name=\"ButtonCreateButton\" class=\"button-wrapper\"><a href=\"https:\/\/www.interconnects.ai\/p\/glm-52-is-the-step-change-for-open?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\" class=\"button primary\" rel=\"nofollow noopener\" target=\"_blank\">Share<\/a><\/p>\n<p>GLM-5.2, in a common naming convention across the industry, looked potentially like an incremental update following the popular GLM-5.1 model. At this point, Moonshot AI, makers of the Kimi models, and Z.ai, makers of the GLM models, have consolidated the top of the reputational market with the most beloved open-weight models among AI researchers. What unfolded is a common lesson in tracking AI models that often minor version numbers can have AI models crossing meaningful user experience thresholds. A small change in benchmarks and training can open a wide range of new use-cases.<\/p>\n<p>What has followed is a slow, groundswell of hype for GLM-5.2. The official, MIT-licensed <a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.2\" rel=\"nofollow noopener\" target=\"_blank\">model weights<\/a> and <a href=\"https:\/\/z.ai\/blog\/glm-5.2\" rel=\"nofollow\">release blog<\/a> dropped three days after the initial rollout, on June 16th. One could ramble many technical details, such as the strong benchmark scores, the very popular RL framework that Z.ai uses (<a href=\"https:\/\/github.com\/THUDM\/slime\" rel=\"nofollow noopener\" target=\"_blank\">SLIME<\/a>), the recommendation of always using the model on Max thinking effort, and so on, but the initial release blogs usually aren\u2019t the thing to focus on. You can wait and read the ecosystem reaction to know if it\u2019s the real deal. <a href=\"https:\/\/www.interconnects.ai\/p\/opus-46-vs-codex-53\" rel=\"nofollow noopener\" target=\"_blank\">Benchmarks are half dead these days<\/a>, anyways.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!xhhJ!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/06\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7074458b-82aa-4658-95bb-9315549abb7f_4239.jpeg\" width=\"1456\" height=\"961\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7074458b-82aa-4658-95bb-9315549abb7f_4239x2799.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:961,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g\" title=\"img_v3_0212o_51684a16-c33f-4429-aea5-9f5f7cdfc30g\"   fetchpriority=\"high\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>What followed on the 16th was a slew of community benchmarks showing better-than-expected results for GLM-5.2. <a href=\"https:\/\/x.com\/arena\/status\/2066943450914943025\" rel=\"nofollow\">Arena\u2019s agent leaderboard<\/a> had it as the only open model mixing it up with OpenAI and Anthropic\u2019s latest models (notably matching Opus 4.8\u2019s no-thinking effort to GLM-5.2\u2019s max mode). This is one of many evals GLM-5.2 is crushing Gemini on, but that\u2019s a topic for another time. A benchmark that has mixed perception in the community (particularly among actual designers), <a href=\"https:\/\/x.com\/Designarena\/status\/2066940737011560652\" rel=\"nofollow\">Design Arena<\/a> even had GLM-5.2 besting Claude Fable itself \u2014 the recently banned hype machine!<\/p>\n<p>Pretty much everyone I respect among the AI commentariat and researcher class has praised the model after using it personally. Such a focal point of discussion among the community has only been so clear with an open model release once before \u2014 <a href=\"https:\/\/www.interconnects.ai\/p\/deepseek-r1-recipe-for-o1\" rel=\"nofollow noopener\" target=\"_blank\">DeepSeek R1<\/a>. This is not a comparison I make lightly, and when I compared <a href=\"https:\/\/www.interconnects.ai\/p\/kimi-k2-and-when-deepseek-moments\" rel=\"nofollow noopener\" target=\"_blank\">Kimi K2\u2019s release to a \u201cDeepSeek Moment,\u201d<\/a> GLM-5.2 has well exceeded that. What made Kimi K2 impressive was that big steps in open model performance could seemingly come from anywhere in China. The step that GLM-5.2 has taken is more of a one way door for AI progress.<\/p>\n<p>Anthropic\u2019s record revenue growth rate on the back of Claude Code is heavily driven by being the best model, and the only model that can really do this. GLM-5.2 is the first of many (coming soon) open weight models to offer credible alternatives. The parallel is very clear, to when DeepSeek R1 showed that open-weight labs, with far fewer resources, could also replicate the chain-of-thought reasoning models that OpenAI championed with o1. As AI systems get more complex and far more expensive to build, with tools, integrated harnesses, and scaled model weights, it was not a given that this GLM-5.2 moment would happen at all.<\/p>\n<p>The key point is that GLM-5.2 is the open weight model that <a href=\"https:\/\/www.interconnects.ai\/p\/claude-code-hits-different?utm_source=publication-search\" rel=\"nofollow noopener\" target=\"_blank\">feels right<\/a> in coding harnesses as a general agent. It\u2019s the first one. I was personally overdue in trying some of the recent peer models, such as Kimi K2.7 or GLM-5.1, but the hype was too much for me to ignore. I put it to work helping make content for my <a href=\"https:\/\/github.com\/natolambert\/rlhf-book\/pull\/457\" rel=\"nofollow noopener\" target=\"_blank\">post-training course<\/a> with Fireworks\u2019 API in Claude Code (<a href=\"https:\/\/docs.fireworks.ai\/ecosystem\/fireconnect\/claude-code\" rel=\"nofollow noopener\" target=\"_blank\">setting this up<\/a> was very easy). There were some minor knife cuts, such as the Claude Code harness \/ my repo documentation trying to send images to the model, which would brick Fireworks API for the session \u2014 forcing a manual context clear. Overall, the model capabilities immediately felt right, and I still have some tinkering to do in which harness and inference provider to use. <\/p>\n<p>For more hype, you can sample the Z.ai founder telling Elon that \u201c<a href=\"https:\/\/x.com\/pmarca\/status\/2067640859957539104\" rel=\"nofollow\">open-weight Fable capabilities will be here sooner than Q1 2027<\/a>,\u201d the CEO of Vercel <a href=\"https:\/\/x.com\/rauchg\/status\/2068517095818809770\" rel=\"nofollow\">saying<\/a> \u201cGenuinely impressed, almost shocked, at how good GLM-5.2 by @zai_org is at coding. This changes things,\u201d and much <a href=\"https:\/\/x.com\/ArtificialAnlys\/status\/2067135640249209175\" rel=\"nofollow\">more<\/a> from a mix of people whose opinions I <a href=\"https:\/\/x.com\/gneubig\/status\/2067936197888930263?s=20\" rel=\"nofollow\">deeply<\/a> <a href=\"https:\/\/x.com\/_xjdr\/status\/2068422921249529916\" rel=\"nofollow\">respect<\/a> and others I\u2019m <a href=\"https:\/\/x.com\/matvelloso\/status\/2067791546335019439?s=20\" rel=\"nofollow\">new to<\/a>.<\/p>\n<p>So, this is a good model, where does this leave us?<\/p>\n<p>There are many trends at play. To start, let\u2019s ground things in the open-closed capabilities gap. I\u2019ve written how I expect an \u201c<a href=\"https:\/\/www.interconnects.ai\/p\/some-ideas-for-what-comes-next-may\" rel=\"nofollow noopener\" target=\"_blank\">explosion in usage<\/a>\u201d if open models crossed the Opus 4.5 in Claude Code threshold from around the start of 2026. Here we are. With Claude Opus 4.5\u2019s release on November 24th, 2025, the gap in time to GLM-5.2\u2019s release on June 16th, 2026 is 204 days \u2014 or about 6.8 months. This puts us square in the 6-9 month time gap that many people claim as the performance lag between the U.S.\u2019s closed labs and China\u2019s open counterparts.<\/p>\n<p>Upon writing this, I\u2019m surprised. As the U.S. labs have so rapidly ramped compute in the last ~year, I\u2019ve expected the gap in performance to grow in time. A very meaningful step in this trajectory will also be Claude Fable 5\u2019s release \u2014 which was more reliant on scale, and therefore the most advanced GPUs, relative to the Claude Opus models. Still, that\u2019s not a satisfactory answer. Continuing to unpack the trajectory here involves more nuance than I can afford to fit in a signposting article.<\/p>\n<p>The most immediate meaning of this is far more serious pricing pressure within the organizations tokenmaxxing, sending Anthropic\u2019s revenue to the moon. Some would predict Anthropic doesn\u2019t realize its forecasted ARR numbers, but I don\u2019t think that prices in the true demand for these models and the inevitable growth. This model existing is a huge boon for the open model economy. All the likes of Fireworks, Together, Thinky (via Tinker), Prime Intellect, and whoever else sells open model inference or finetuning just hit another inflection point. <\/p>\n<p>It\u2019ll take a long time for the effects here to diffuse into the broader economy (and use-cases). Workflows are becoming more complex, with people using different models for planning, primary coding, and subagent dispatch. I expect the hype to continue to grow, and heck, as I\u2019m writing this on a Sunday evening, I could see the media and market reaction on the Monday being a thing just like the DeepSeek R1 release. This diffusion happening while Anthropic\u2019s, and by extension the U.S.\u2019s flagship model, is still banned is a severe economic dagger. GLM-5.2 is being given time to carve out the economic underbelly of the frontier labs when they want to be pushing forward into higher margin, higher revenue domains enabled only by the absolute frontier models.<\/p>\n<p>The economic concern mirrors a story that has been told many times in AI, so it\u2019s unclear when it\u2019ll stick.<\/p>\n<p>The conversation that feels more core to the trajectory of AI is that of regulation and control of open models. I think it is an economic good for cheap intelligence to diffuse widely, and our default position should be to cheer for open models, but this model\u2019s release date will have it be permanently associated with Claude Fable \u2014 and therefore Claude Mythos \u2014 in the mental map of AI power structures. We are at a point where Mythos-class model capabilities are deemed not safe for release by the U.S. Government and the Chinese model makers are charging forward in capabilities available to all. <\/p>\n<p>These trend lines aren\u2019t necessarily causally linked, as we don\u2019t know the cyber performance of GLM-5.2 versus its predecessors, but the capabilities are definitely correlated. Without anything changing, this points to a potentiality where the U.S. Government decides a certain open-weights Chinese model is not safe for the public. There are many other potential scenarios here too, but what is clear is that we have a lot of work to do in mapping them out, preparing our infrastructure, and messaging to society. <\/p>\n<p>It\u2019ll take a lot more people than just me to imagine and communicate a world to decision makers for how to manage evermore capable open models. We have years more of AI progress to come, with Nvidia\u2019s next generation chips already in production and a constant stream of algorithmic advancements. It feels like a narrow path for open model advocates to take, but we need to figure out how to make them viable so the massive leaps in performance don\u2019t only go to closed models. <\/p>\n<p>I totally see why it is scary to imagine an openly accessible Mythos class model, but if open models get banned now and only closed models get 10 or 100X better in 2 years in the hands of one or two companies, I think we will have bigger problems on our hands.<\/p>\n","protected":false},"excerpt":{"rendered":"Housekeeping: Following my \u201cState of the blog\u201d post last week, noting a slight increase in paid features, it\u2019s&hellip;\n","protected":false},"author":2,"featured_media":652810,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-652809","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/652809","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=652809"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/652809\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/652810"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=652809"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=652809"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=652809"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}