{"id":791298,"date":"2026-07-10T12:08:23","date_gmt":"2026-07-10T12:08:23","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/791298\/"},"modified":"2026-07-10T12:08:23","modified_gmt":"2026-07-10T12:08:23","slug":"grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/791298\/","title":{"rendered":"Grok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter much"},"content":{"rendered":"<p>Update, June 9, 2026:<\/p>\n<p>Artificial Analysis ranks Grok 4.5 fourth on its Intelligence Index, behind Fable 5, GPT-5.5, and Opus 4.8. The model gained 16 points over Grok 4.3, putting SpaceXAI close to frontier performance. It trails OpenAI and Anthropic but beats all open-weights models and Google&#8217;s Gemini.<\/p>\n<p>Grok 4.5 is also very cost-efficient on the Intelligence Index, where a single task costs just $0.31. That&#8217;s less than GLM-5.2 and Kimi K2.6, and five times cheaper than Claude Sonnet 5 (max), which scores lower on the same index. The independent AI benchmark service measures model performance across a range of tasks and aggregates the results into a single score.<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/07\/grok_45_benchmark-1-scaled.jpg\"><img fetchpriority=\"high\" decoding=\"async\" class=\"size-full wp-image-37540\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/07\/grok_45_benchmark-1-scaled.jpg\" alt=\"\" width=\"2560\" height=\"2057\"\/><\/a>Grok 4.5 is a big leap for xAI, even if it doesn&#8217;t reach the top. And it&#8217;s cheap. | Image: Artificial Analysis<\/p>\n<p><a href=\"https:\/\/x.com\/ArtificialAnlys\/status\/2074956932289282087\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">According to Artificial Analysis<\/a>, Grok 4.5 performs particularly well on agentic tasks. On the Coding Agent Index, Grok 4.5 running in Grok Build, xAI&#8217;s equivalent to Claude Code, scores 76 points, matching GPT-5.5 in Codex and trailing Fable 5 in Claude Code by just one point, at a fraction of the cost. Per task, Grok 4.5 in Grok Build costs $2.49, compared to $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code. Grok 4.5 also averages just 1.9 million tokens per task, far less than GPT-5.5 (6.2M) and Fable 5 (7.2M).<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/07\/AA_cost_per_task.png\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-37541 size-full\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/07\/AA_cost_per_task.png\" alt=\"\" width=\"1478\" height=\"1155\"\/><\/a>In its performance tier, Grok 4.5 is very cheap and efficient, assuming benchmark results hold up. | Image: Artificial Analysis<\/p>\n<p><a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/x.com\/ArtificialAnlys\/status\/2074956952271225248\">Artificial Analysis also flags a weakness<\/a>. Accuracy on the AA-Omniscience Index rose from 35 to 52 percent, but the hallucination rate jumped from 25 to 54 percent too. The model knows more, but it&#8217;s also more confident when it&#8217;s wrong.<\/p>\n<p>Original article, June 8, 2026:<\/p>\n<p>xAI has released Grok 4.5. The model was trained on tens of thousands of Nvidia GB300 GPUs and targets coding, agentic tasks, and knowledge work.<\/p>\n<p>Benchmark results paint a mixed picture. On Terminal Bench 2.1, which tests complex command-line tasks, Grok 4.5 scores 83.3%, nearly matching GPT 5.5 (83.4%) and trailing Anthropic&#8217;s\u00a0<a href=\"https:\/\/the-decoder.com\/anthropics-fable-5-is-back-worldwide-after-a-two-week-government-ban-over-a-jailbreak\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Fable 5<\/a>\u00a0(84.3%) by just one point.<\/p>\n<p>But the gaps widen elsewhere. On DeepSWE 1.1, which measures the ability to resolve real GitHub issues, Grok 4.5 hits 53%, well behind OpenAI&#8217;s\u00a0<a href=\"https:\/\/the-decoder.com\/gpt-5-5-costs-49-to-92-percent-more-than-its-predecessor-depending-on-the-input-length\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">GPT-5.5<\/a>\u00a0at 67% and Fable 5 at 70%. On SWE Bench Pro, a curated set of harder software engineering problems, it scores 64.7%, beating Opus 4.8 (69.2% with max settings) in some configurations but falling short of Fable 5&#8217;s 80.4%.<\/p>\n<p>Model<br \/>\nDeepSWE 1.1<br \/>\nTerminal Bench 2.1<br \/>\nSWE Bench Pro<\/p>\n<p>Fable max<br \/>\n70%<br \/>\n84.3%<br \/>\n80.4%<\/p>\n<p>GPT 5.5 xhigh<br \/>\n67%<br \/>\n83.4%<br \/>\n58.6%<\/p>\n<p>Opus 4.8 max<br \/>\n59%<br \/>\n78.9%<br \/>\n69.2%<\/p>\n<p>Grok 4.5<br \/>\n53%<br \/>\n83.3%<br \/>\n64.7%<\/p>\n<p>GLM 5.2<br \/>\n44%<br \/>\n81.0%<br \/>\n62.1%<\/p>\n<p><a target=\"_blank\" rel=\"noopener nofollow\" href=\"https:\/\/x.ai\/news\/grok-4-5\">xAI says<\/a> it relied on heavy data filtering, deduplication, and domain-specific selection during training to keep data quality high. The reinforcement learning stage covered hundreds of thousands of tasks, mostly from software engineering, with automated scoring. xAI built the training infrastructure for asynchronous learning, so agentic runs could stretch over many hours while training continued in parallel.<\/p>\n<p>Grok 4.5 undercuts the competition on price<\/p>\n<p>Grok 4.5 costs $2 per million input tokens and $6 per million output tokens. That&#8217;s already far below the competition. Opus 4.8 runs $5 input and $25 output per million tokens. Fable 5 charges $10 input and $50 output per million tokens. GPT-5.5 and\u00a0<a href=\"https:\/\/the-decoder.de\/openai-stellt-gpt-5-6-sol-vor-und-uebertrifft-claude-mythos-in-zentralen-benchmarks\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">GPT-5.6 sit at $5 input and $30 output<\/a>.<\/p>\n<p>xAI also says Grok 4.5 uses 4.2 times fewer tokens than\u00a0<a href=\"https:\/\/the-decoder.com\/anthropic-ships-claude-opus-4-8-as-a-modest-but-tangible-improvement-that-tops-gpt-5-5-in-most-benchmarks\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Opus 4.8<\/a>\u00a0on SWE Bench Pro tasks and delivers results at 80 tokens per second.\u00a0<a href=\"https:\/\/the-decoder.com\/frontier-radar-3-how-agentic-ai-is-turning-tokens-into-a-business-metric\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Lower per-token pricing and fewer tokens per task<\/a>\u00a0make Grok 4.5 by far the cheapest option in this performance tier, assuming the performance and efficiency gains hold up in practice.<\/p>\n<p>Model<br \/>\nInput (per 1M tokens)<br \/>\nOutput (per 1M tokens)<\/p>\n<p>Grok 4.5<br \/>\n$2<br \/>\n$6<\/p>\n<p>Opus 4.8<br \/>\n$5<br \/>\n$25<\/p>\n<p>GPT-5.5 \/ GPT-5.6<br \/>\n$5<br \/>\n$30<\/p>\n<p>Fable 5<br \/>\n$10<br \/>\n$50<\/p>\n<p>The pricing strategy echoes what Chinese vendors like <a href=\"https:\/\/the-decoder.com\/zhipu-ais-glm-5-2-closes-in-on-closed-source-leaders-in-coding-marathons\/\" rel=\"nofollow noopener\" target=\"_blank\">Zhipu<\/a> and <a href=\"https:\/\/the-decoder.com\/as-agentic-ai-pushes-rivals-to-raise-prices-and-cap-usage-deepseek-ships-a-good-enough-model-for-almost-nothing\/\" rel=\"nofollow noopener\" target=\"_blank\">DeepSeek<\/a> have been doing: get close enough on performance, then win on price.<\/p>\n<p>Grok 4.5 is available now through Grok Build,\u00a0<a href=\"https:\/\/cursor.com\/blog\/spacex-model-training\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Cursor<\/a>, and the\u00a0<a href=\"https:\/\/console.x.ai\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">xAI console<\/a>. Plugins are live for\u00a0<a href=\"https:\/\/marketplace.microsoft.com\/en-us\/product\/office\/WA200011055?tab=Overview\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Word<\/a>,\u00a0<a href=\"https:\/\/marketplace.microsoft.com\/en-us\/product\/office\/WA200011057?tab=Overview\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">PowerPoint<\/a>, and\u00a0<a href=\"https:\/\/marketplace.microsoft.com\/en-us\/product\/office\/WA200011056?tab=Overview\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">Excel<\/a>. The model isn&#8217;t available in the EU yet, with xAI targeting a mid-July launch. xAI trained Grok 4.5 alongside the code editor Cursor, which\u00a0<a href=\"https:\/\/the-decoder.com\/spacex-bets-60-billion-on-cursor-to-catch-openai-and-anthropic\/\" target=\"_blank\" rel=\"noopener noreferrer nofollow\">SpaceX acquired in mid-June for $60 billion in stock<\/a>.<\/p>\n<p>\t\t\t\tAI News Without the Hype \u2013 Curated by Humans<\/p>\n<p>\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive &#8220;AI Radar&#8221; frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t<\/p>\n<p>\t\t\t\t<a href=\"https:\/\/the-decoder.com\/subscription\/\" class=\"inline-block text-white bg-(--heise-primary) mt-3 hover:bg-blue-800 focus:ring-4 focus:outline-none focus:ring-blue-300 font-medium rounded-sm w-full sm:w-auto  pl-3 pr-3 py-2.5 text-center newsletter-submit-button hover:no-underline\" rel=\"nofollow noopener\" target=\"_blank\"><br \/>\n\t\t\t\t\tSubscribe now\t\t\t\t<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Update, June 9, 2026: Artificial Analysis ranks Grok 4.5 fourth on its Intelligence Index, behind Fable 5, GPT-5.5,&hellip;\n","protected":false},"author":2,"featured_media":791299,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[256,254,255,64,63,10646,105,3123],"class_list":["post-791298","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-au","tag-australia","tag-grok","tag-technology","tag-xai"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/791298","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=791298"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/791298\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/791299"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=791298"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=791298"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=791298"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}