{"id":445459,"date":"2026-05-16T23:44:11","date_gmt":"2026-05-16T23:44:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/445459\/"},"modified":"2026-05-16T23:44:11","modified_gmt":"2026-05-16T23:44:11","slug":"latest-open-artifacts-21-open-model-bonanza-gemma-4-deepseek-v4-kimi-k2-6-mimo-2-5-glm-5-1-others","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/445459\/","title":{"rendered":"Latest open artifacts (#21): Open model bonanza! Gemma 4, DeepSeek V4, Kimi K2.6, MiMo 2.5, GLM-5.1 &#038; others"},"content":{"rendered":"<p>This month was packed, with all open frontier labs, including DeepSeek, releasing new models. The latter prompted an evaluation by the <a href=\"https:\/\/www.nist.gov\/news-events\/news\/2026\/05\/caisi-evaluation-deepseek-v4-pro\" rel=\"nofollow noopener\" target=\"_blank\">Center for AI Standards and Innovation (CAISI)<\/a>, which has evaluated open models and their risks in the past. Their result is that open models lag behind the American frontier, with the gap becoming wider over time:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!4DPW!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/9d954e79-7538-48b7-bd3c-c6cd21421329_2800.jpeg\" width=\"1456\" height=\"965\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/9d954e79-7538-48b7-bd3c-c6cd21421329_2800x1856.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:965,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains.\" title=\"Comparison of aggregate capabilities over time of the most capable publicly released U.S. and PRC models according to a suite of benchmarks covering five domains.\"   fetchpriority=\"high\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>For the report, they calculate an Elo score based on <a href=\"https:\/\/en.wikipedia.org\/wiki\/Item_response_theory\" rel=\"nofollow noopener\" target=\"_blank\">Item Response Theory<\/a>, which is commonly used to compare different models, even when they were tested on a different set of benchmarks. For V4, CAISI used nine different benchmarks:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!cVg4!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/206a5c61-cc5b-423e-8db6-28f5f209291c_1546.jpeg\" width=\"1456\" height=\"1209\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1209,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:215444,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.interconnects.ai\/i\/197676648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F206a5c61-cc5b-423e-8db6-28f5f209291c_1546x1284.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   class=\"sizing-normal\"\/><\/a><\/p>\n<p>The huge Elo difference is explained by DeepSeek V4s bad score in CTF-Archive-Diamond (which was only run with a subset of the benchmark and extrapolated with IRT for V4), PortBench (a CAISI-private benchmark) and ARC-AGI-2 (with a different scoring method than the public leaderboards). The differences in these benchmark have a huge impact on the overall Elo, which can exacerbate the difference in capabilities. <\/p>\n<p>When using <a href=\"https:\/\/epoch.ai\/eci\" rel=\"nofollow noopener\" target=\"_blank\">Epoch AI\u2019s ECI<\/a>, which also uses IRT over a set of different benchmarks, we see that the gap roughly stays between 3-7 months since R1:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!qx4F!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400.jpeg\" width=\"1456\" height=\"910\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:910,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2762591,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.interconnects.ai\/i\/197676648?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F608dea3a-c43c-4bf0-a5c2-f5b9bb7e63d3_2400x1500.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   class=\"sizing-normal\"\/><\/a>The open&lt;&gt;closed gap in ECI (from https:\/\/mcnair.center\/china\/)<\/p>\n<p>However, both CAISI and ECI paint an incomplete picture, as both use standardized (and simple) setups to compare the capabilities of models. To be more concrete: Coding tasks are evaluated using access to bash and a for-loop with a fixed budget of tokens, not with a harness such as Claude Code or OpenCode, which models are trained in! These setups result in benchmarks claiming that porting applications to another language is <a href=\"https:\/\/programbench.com\/\" rel=\"nofollow noopener\" target=\"_blank\">currently not possible<\/a>, while <a href=\"https:\/\/github.com\/oven-sh\/bun\/pull\/30412\" rel=\"nofollow noopener\" target=\"_blank\">Bun has been ported from Zig to Rust with 1 million LOC changes<\/a>.<\/p>\n<p>Therefore, we would argue that a frontier comparison of open and closed models would also need to elicit the capabilities of all models better, which means the usage of the preferred harnesses, as well as model-specific prompting.<\/p>\n<p>This section was written primarily by Florian. An interesting dynamic within Interconnects is that Florian believes more in the proximity of open frontier models to closed alternatives in true performance. Nathan thinks the benchmarks are imperfect as well, but thinks the closed models are ahead by more. We\u2019re going to continue to unpack this in our future content.<\/p>\n<p data-attrs=\"{&quot;url&quot;:&quot;https:\/\/www.interconnects.ai\/p\/latest-open-artifacts-21-open-model?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}\" data-component-name=\"ButtonCreateButton\" class=\"button-wrapper\"><a href=\"https:\/\/www.interconnects.ai\/p\/latest-open-artifacts-21-open-model?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\" class=\"button primary\" rel=\"nofollow noopener\" target=\"_blank\">Share<\/a><\/p>\n<p><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.5-Pro\" rel=\"nofollow noopener\" target=\"_blank\">MiMo-V2.5-Pro<\/a> by <a href=\"https:\/\/huggingface.co\/XiaomiMiMo\" rel=\"nofollow noopener\" target=\"_blank\">XiaomiMiMo<\/a>: Avid Artifacts readers know that Xiaomi has been working on open models for a while; its debut was <a href=\"https:\/\/www.interconnects.ai\/p\/latest-open-artifacts-10-new-deepseek\" rel=\"nofollow noopener\" target=\"_blank\">exactly one year ago<\/a>. The progress of its releases is remarkable, with 2.5 Pro (released under Apache 2.0) being neck and neck with other flagship models such as Kimi K2.6 and GLM-5.1 in both benchmarks and <a href=\"https:\/\/x.com\/Designarena\/status\/2054776484833952000?s=20\" rel=\"nofollow\">real-world usage<\/a>.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!25fp!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/34228788-0da1-4aea-840e-5eae68bbfa17_1200.jpeg\" width=\"1200\" height=\"786\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/34228788-0da1-4aea-840e-5eae68bbfa17_1200x786.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:786,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"Image\" title=\"Image\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p><a href=\"https:\/\/huggingface.co\/google\/gemma-4-26B-A4B-it\" rel=\"nofollow noopener\" target=\"_blank\">gemma-4-26B-A4B-it<\/a> by <a href=\"https:\/\/huggingface.co\/google\" rel=\"nofollow noopener\" target=\"_blank\">google<\/a> (full Interconnects post <a href=\"https:\/\/www.interconnects.ai\/p\/gemma-4-and-what-makes-an-open-model\" rel=\"nofollow noopener\" target=\"_blank\">here<\/a>): The long-awaited update to the Gemma series, featuring multiple sizes: <a href=\"https:\/\/huggingface.co\/google\/gemma-4-E2B-it\" rel=\"nofollow noopener\" target=\"_blank\">4B<\/a>, <a href=\"https:\/\/huggingface.co\/google\/gemma-4-E4B-it\" rel=\"nofollow noopener\" target=\"_blank\">9B<\/a>, and <a href=\"https:\/\/huggingface.co\/google\/gemma-4-31B-it\" rel=\"nofollow noopener\" target=\"_blank\">31B<\/a> dense models, as well as a 26B-A4B MoE. Even more importantly, with Gemma 4, Google has decided to use Apache 2.0 as its license, which removes the uncertainty and legal challenges around interpreting custom licenses.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/moonshotai\/Kimi-K2.6\" rel=\"nofollow noopener\" target=\"_blank\">Kimi-K2.6<\/a> by <a href=\"https:\/\/huggingface.co\/moonshotai\" rel=\"nofollow noopener\" target=\"_blank\">moonshotai<\/a>: An update to the Kimi series, delivering stronger performance across the board and making it one of the best open models out there yet again. They also focus on long-horizon performance, showing that open models are capable of running over hours to complete tasks or optimize performance. Given the focus of everyone to build <a href=\"https:\/\/github.com\/karpathy\/autoresearch\" rel=\"nofollow noopener\" target=\"_blank\">autoresearch<\/a>-like systems, seeing open models catch up is important.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!K_mJ!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/05\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/e08af2e3-9ce3-4ace-bec7-e8aecce5e120_1018.jpeg\" width=\"1456\" height=\"932\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/e08af2e3-9ce3-4ace-bec7-e8aecce5e120_10188x6520.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:932,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;K2.6 Qwen3.5-0.8B Mac inference optimization case&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"K2.6 Qwen3.5-0.8B Mac inference optimization case\" title=\"K2.6 Qwen3.5-0.8B Mac inference optimization case\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p><a href=\"https:\/\/huggingface.co\/poolside\/Laguna-XS.2\" rel=\"nofollow noopener\" target=\"_blank\">Laguna-XS.2<\/a> by <a href=\"https:\/\/huggingface.co\/poolside\" rel=\"nofollow noopener\" target=\"_blank\">poolside<\/a>: Poolside AI has released its first public coding-focused models, including the open-weight XS.2. Its size (33B-A3B) makes it attractive for local use, with performance on par with other models in that size range. The accompanying <a href=\"https:\/\/poolside.ai\/blog\/laguna-a-deeper-dive\" rel=\"nofollow noopener\" target=\"_blank\">blog post<\/a> is worth a read, as is <a href=\"https:\/\/poolside.ai\/blog\/through-the-looking-glass\" rel=\"nofollow noopener\" target=\"_blank\">the deep dive<\/a> into reward hacking during coding evaluations.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Flash\" rel=\"nofollow noopener\" target=\"_blank\">DeepSeek-V4-Flash<\/a> by <a href=\"https:\/\/huggingface.co\/deepseek-ai\" rel=\"nofollow noopener\" target=\"_blank\">deepseek-ai<\/a>: DeepSeek has finally released its successor to the V3 series, which it kept updating for months. It comes in two sizes: Pro, which is a 1.6T-A49B MoE, and Flash, a 284B-13B model. Based on others\u2019 experience, the latter model seems to be the real star of the show, as its performance is relatively strong, while Pro seems to underdeliver relative to its size. The <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Pro\/blob\/main\/DeepSeek_V4.pdf\" rel=\"nofollow noopener\" target=\"_blank\">tech report<\/a> goes into great detail, including the architectural changes used to achieve better and cheaper long-context performance.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.6-35B-A3B\" rel=\"nofollow noopener\" target=\"_blank\">Qwen3.6-35B-A3B<\/a> by <a href=\"https:\/\/huggingface.co\/Qwen\" rel=\"nofollow noopener\" target=\"_blank\">Qwen<\/a>: An update to the Qwen 3.5 series targeting one of the most widely used sizes.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M\" rel=\"nofollow noopener\" target=\"_blank\">LFM2.5-350M<\/a> by <a href=\"https:\/\/huggingface.co\/LiquidAI\" rel=\"nofollow noopener\" target=\"_blank\">LiquidAI<\/a>: With 28T tokens for 350M parameters, this model might be the most overtrained model out there.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/arcee-ai\/Trinity-Large-Thinking\" rel=\"nofollow noopener\" target=\"_blank\">Trinity-Large-Thinking<\/a> by <a href=\"https:\/\/huggingface.co\/arcee-ai\" rel=\"nofollow noopener\" target=\"_blank\">arcee-ai<\/a>: The reasoning version of Trinity, one of the best Western open models. It has topped the OpenRouter charts for a while and can power agentic applications such as OpenClaw.<\/p>\n<p><a href=\"https:\/\/huggingface.co\/zai-org\/GLM-5.1\" rel=\"nofollow noopener\" target=\"_blank\">GLM-5.1<\/a> by <a href=\"https:\/\/huggingface.co\/zai-org\" rel=\"nofollow noopener\" target=\"_blank\">zai-org<\/a>: An update to GLM-5, improving scores across the board. The focus for this update is on long-horizon tasks.<\/p>\n","protected":false},"excerpt":{"rendered":"This month was packed, with all open frontier labs, including DeepSeek, releasing new models. The latter prompted an&hellip;\n","protected":false},"author":2,"featured_media":445460,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,343,344,85,46,125],"class_list":["post-445459","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-il","tag-israel","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/445459","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=445459"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/445459\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/445460"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=445459"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=445459"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=445459"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}