{"id":566651,"date":"2026-07-30T09:56:08","date_gmt":"2026-07-30T09:56:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/566651\/"},"modified":"2026-07-30T09:56:08","modified_gmt":"2026-07-30T09:56:08","slug":"openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-with-its-latest-api-and-two-additional-settings","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/566651\/","title":{"rendered":"OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings"},"content":{"rendered":"<p>Update:<\/p>\n<p>ARC Prize co-founder\u00a0<a href=\"https:\/\/x.com\/fchollet\/status\/2082732210436575669\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Fran\u00e7ois Chollet responded to OpenAI&#8217;s results<\/a>\u00a0by distinguishing between two kinds of test setups. Harnesses &#8220;custom-made to solve the benchmark or that contain knowledge about the benchmark format&#8221; are off limits, he said. General-purpose API settings &#8220;that were not developed for ARC-AGI-3 and that are available to all API users&#8221; are fair game. In effect, Chollet is conceding that ARC Prize&#8217;s own GPT-5.6 Sol score put OpenAI at a disadvantage.<\/p>\n<p>He noted that ARC Prize has had &#8220;a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction,&#8221; and welcomed the company &#8220;starting to figure out the answer.&#8221; Different providers using different settings does create &#8220;a potential parity issue,&#8221; Chollet said, but he considers that acceptable &#8220;as long as the settings and the cost are clearly reported.&#8221;<\/p>\n<p>Original article:<\/p>\n<p>OpenAI says it can keep up on ARC-AGI-3.\u00a0After Anthropic&#8217;s\u00a0<a href=\"https:\/\/the-decoder.com\/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Claude Opus 5 quadrupled the record score on the logic benchmark<\/a>,\u00a0<a href=\"https:\/\/openai.com\/index\/how-two-settings-tripled-our-arc-agi-3-scores\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">OpenAI is now showing<\/a>\u00a0that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5&#8217;s 30.2 percent.<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/GPT-5.6-Sol-on-the-ARC-AGI-3-Public-Set.png\"><img fetchpriority=\"high\" decoding=\"async\" class=\"wp-image-38356 size-full\" src=\"https:\/\/www.newsbeep.com\/il\/wp-content\/uploads\/2026\/07\/GPT-5.6-Sol-on-the-ARC-AGI-3-Public-Set.png\" alt=\"\" width=\"1290\" height=\"898\"\/><\/a>GPT-5.6 Sol&#8217;s ARC-AGI-3 scores jump dramatically when using OpenAI&#8217;s custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI<\/p>\n<p>OpenAI isn&#8217;t using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own <a href=\"https:\/\/the-decoder.com\/openai-upgrades-responses-api-with-features-built-specifically-for-long-running-ai-agents\/\" rel=\"nofollow noopener\" target=\"_blank\">Responses API<\/a> with &#8220;Retained Reasoning,&#8221; which keeps the model&#8217;s chain of thought between steps, and\u00a0&#8220;Compaction,&#8221;\u00a0which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model&#8217;s reasoning gets discarded after each action.<\/p>\n<p>OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That&#8217;s true, and\u00a0<a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/arxiv.org\/pdf\/2603.24621\">ARC-AGI-3\u00a0is designed<\/a> to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons,\u00a0<a href=\"https:\/\/x.com\/arcprize\/status\/2082672003765670160\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">ARC Prize said in response to OpenAI&#8217;s results<\/a>. The sticking point is whether ARC Prize used an older &#8220;OpenAI-style completions API&#8221; that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Update: ARC Prize co-founder\u00a0Fran\u00e7ois Chollet responded to OpenAI&#8217;s results\u00a0by distinguishing between two kinds of test setups. Harnesses &#8220;custom-made&hellip;\n","protected":false},"author":2,"featured_media":566652,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[1463,85,46,1748,125],"class_list":["post-566651","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-anthropic","tag-il","tag-israel","tag-openai","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/566651","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=566651"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/566651\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/566652"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=566651"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=566651"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=566651"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}