{"id":785426,"date":"2026-09-28T16:08:38","date_gmt":"2026-09-28T16:08:38","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/785426\/"},"modified":"2026-09-28T16:08:38","modified_gmt":"2026-09-28T16:08:38","slug":"wheres-the-intelligence-explosion-by-ramez-naam","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/785426\/","title":{"rendered":"Where\u2019s the \u201cintelligence explosion\u201d? &#8211; by Ramez Naam"},"content":{"rendered":"<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!x5a8!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ec0b36-c40e-47a2-831f-16108c969d77_1075x717.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/08ec0b36-c40e-47a2-831f-16108c969d77_1075.jpeg\" width=\"718\" height=\"478.8893023255814\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/08ec0b36-c40e-47a2-831f-16108c969d77_1075x717.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:717,&quot;width&quot;:1075,&quot;resizeWidth&quot;:718,&quot;bytes&quot;:191581,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https:\/\/www.noahpinion.blog\/i\/217632453?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ec0b36-c40e-47a2-831f-16108c969d77_1075x717.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   fetchpriority=\"high\" class=\"sizing-normal\"\/><\/a>Art by GPT-6<\/p>\n<p>One of my fundamental beliefs about the world is that <a href=\"https:\/\/x.com\/ramez\" rel=\"nofollow\">Ramez Naam<\/a> ought to blog more. Ramez is one of the world\u2019s greatest futurists \u2014 he <a href=\"https:\/\/www.scientificamerican.com\/blog\/guest-blog\/smaller-cheaper-faster-does-moores-law-apply-to-solar-cells\/\" rel=\"nofollow noopener\" target=\"_blank\">predicted the solar<\/a> and battery revolutions long before these were widely understood. If you were reading Ramez in 2011, you were able to understand the future of both energy technology and climate change, long before other people did. His earlier book More than Human is still a great guide to the kind of biological enhancements that AI might make possible. Ramez is also an excellent science fiction author, having written <a href=\"https:\/\/www.amazon.com\/dp\/0857662929?lv=shuf&amp;channelId=500&amp;plpRedirect=mhFallback\" rel=\"nofollow noopener\" target=\"_blank\">a trilogy of novels<\/a> in which nanotechnological telepathy is distributed as a party drug (I\u2019m not sure if he actually expects that to happen, but it\u2019s a very cool idea). <\/p>\n<p>Unfortunately, although <a href=\"https:\/\/www.rameznaam.com\/\" rel=\"nofollow noopener\" target=\"_blank\">he does have a Substack<\/a> (which you should absolutely follow), Ramez does not blog regularly. However, after having a lengthy private debate with him about Recursive Self-Improvement, I was able to prevail upon him to write up his thoughts for my blog. <\/p>\n<p>To say that RSI is a big deal in the AI world would be a colossal understatement. Among AI researchers, entrepreneurs, and AI safety people, there\u2019s a widespread belief that as AI gets better at improving itself, there will be a \u201cfast takeoff\u201d or \u201cFOOM\u201d, in which AI\u2019s capabilities \u201ctake off\u201d and create a <a href=\"https:\/\/en.wikipedia.org\/wiki\/Technological_singularity\" rel=\"nofollow noopener\" target=\"_blank\">technological Singularity<\/a>. This event is a staple of science fiction, including works by my favorite sci-fi author, <a href=\"https:\/\/www.noahpinion.blog\/p\/go-read-some-vernor-vinge\" rel=\"nofollow noopener\" target=\"_blank\">Vernor Vinge<\/a>. <\/p>\n<p>A lot of people in the industry believe that this moment is now close at hand, and are racing toward that prize:<\/p>\n<p><a href=\"https:\/\/x.com\/tanayj\/status\/2104004090007302287\" target=\"_blank\" rel=\"noopener noreferrer nofollow\" data-component-name=\"Twitter2ToDOM\" class=\"pencraft pc-display-contents pc-reset\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/pbs.substack.com\/profile_images\/806762286383665156\/peWa7SS9.jpg\" alt=\"X avatar for @tanayj\"  width=\"40\" height=\"40\" draggable=\"false\" class=\"img-OACg1c object-fit-cover-u4ReeV pencraft pc-reset\"\/><\/p>\n<p>Tanay Jaipuria@tanayj<\/p>\n<p>Noam Brown says that recursive self-improvement \/ the ability of AI models to do AI research is the top priority for OpenAI by a wide margin.  <\/p>\n<p>OpenAI ranks it above making models that sell.<\/p>\n<p>(via @theinformation podcast)<\/p>\n<p>12:24 AM \u00b7 Sep 27, 2026 \u00b7 4.01K Views<\/p>\n<p>4 Replies \u00b7 2 Reposts \u00b7 44 Likes<\/p>\n<p><\/a><\/p>\n<p>But Ramez \u2014 normally among the most wide-eyed of techno-optimists \u2014 is highly skeptical that we\u2019ll see anything like the \u201cFOOM\u201d of Vernor Vinge novels. In this lengthy, well-researched post, he explains his skepticism. <\/p>\n<p>Personally, I\u2019m agnostic. Ramez\u2019s case necessarily rests on a lot of assumptions; although it\u2019s cogently laid out, I think the real answer is that we\u2019ll just have to wait and see whether the Singularity arrives. But even more fundamentally, I don\u2019t know how much this debate matters in the practical sense \u2014 even without the kind of Singularity depicted in sci-fi novels, AI capabilities are improving so rapidly that they\u2019re already superhuman in many respects, and soon will probably be strongly superhuman in most or all dimensions. The AI of 2040 is going to look godlike, whether or not it explodes into an actual god in 2027. <\/p>\n<p>Still, it\u2019s a very interesting argument, and Ramez\u2019s thoughts on the future of technology are always worth listening to. <\/p>\n<p>AI is already helping improve itself. The question is whether even fully autonomous recursive self-improvement (RSI) would cause a runaway intelligence explosion.<\/p>\n<p>The theory is that each generation of AI could build a better successor, faster than the last generation did. That could lead to a \u201cfast takeoff,\u201d with capabilities surging to artificial superintelligence (ASI) in a year, months, or even days.<\/p>\n<p>Here\u2019s my take: Given our best current data, the AI self-improvement loop would need to be roughly 5\u201310\u00d7 stronger to sustain itself, let alone run away. I\u2019ll explain this math in <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A78-the-self-improvement-loop-doesnt-look-strong-enough\" rel=\"nofollow noopener\" target=\"_blank\">section 8<\/a>. I expect incredibly rapid AI progress by the standards of nearly any other technology. But the evidence we have doesn\u2019t suggest a sudden explosion to incomprehensible superintelligence anytime soon.<\/p>\n<p>I could be wrong. <a href=\"https:\/\/forecastingresearch.org\/research\/ai-progress-accuracy-update\" rel=\"nofollow noopener\" target=\"_blank\">Forecasters have repeatedly underestimated AI progress<\/a>! I could well be next. One thing that\u2019s clear is that <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A710-we-need-more-data\" rel=\"nofollow noopener\" target=\"_blank\">we need better data<\/a>. For now, let\u2019s work with what we can measure, and stay open to breakthroughs that could change the picture.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!cT_g!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59d5d33b-af61-4e3c-946e-e14a227f0b04_4000x2374.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/59d5d33b-af61-4e3c-946e-e14a227f0b04_4000.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/59d5d33b-af61-4e3c-946e-e14a227f0b04_4000x2374.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1144311,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F59d5d33b-af61-4e3c-946e-e14a227f0b04_4000x2374.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" fetchpriority=\"high\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 1. How strong is the self-improvement loop? <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Model<\/a>.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7what-the-future-holds\" rel=\"nofollow noopener\" target=\"_blank\">Jump to the conclusion<\/a>.<\/p>\n<p>Here\u2019s the case, with links to each part:<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7we-already-have-narrow-superintelligence\" rel=\"nofollow noopener\" target=\"_blank\">Narrow<\/a><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7we-already-have-narrow-superintelligence\" rel=\"nofollow noopener\" target=\"_blank\"> Superintelligence Is Here Today<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A73-real-ai-research-is-harder-than-benchmarks-or-forecasts\" rel=\"nofollow noopener\" target=\"_blank\">Real-World Research Is Harder<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A74-the-sharp-diminishing-returns-to-impressive-ai-numbers\" rel=\"nofollow noopener\" target=\"_blank\">Impressive AI Numbers \u2192 Sharp Diminishing Returns<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A76-were-not-seeing-runaway-acceleration\" rel=\"nofollow noopener\" target=\"_blank\">We\u2019re Not Seeing Signs of Acceleration<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7keeping-up-the-pace-takes-exponentially-more-resources\" rel=\"nofollow noopener\" target=\"_blank\">Keeping Up the Pace Takes Exponentially More Resources<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7better-ai-may-be-needed-just-to-maintain-the-pace\" rel=\"nofollow noopener\" target=\"_blank\">Better AI May Be Needed Just to Maintain the Pace<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A77-why-does-progress-get-harder\" rel=\"nofollow noopener\" target=\"_blank\">Progress Gets Harder; Ideas Get Harder to Find<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A78-the-self-improvement-loop-doesnt-look-strong-enough\" rel=\"nofollow noopener\" target=\"_blank\">The Current Feedback Loop Doesn\u2019t Look Strong Enough<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7how-fast-are-gains-coming-now\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI\u2019s Data Shows How Weak the Loop Is<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A79-what-could-accelerate-this\" rel=\"nofollow noopener\" target=\"_blank\">What Could Accelerate Progress?<\/a><\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A710-we-need-more-data\" rel=\"nofollow noopener\" target=\"_blank\">We Need More Data to Track This Well<\/a><\/p>\n<p>Key charts: <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7how-strong-is-the-feedback-loop\" rel=\"nofollow noopener\" target=\"_blank\">The feedback loop<\/a> \u00b7 <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7measured-progress-vs-ai-2027-and-eci-extrapolated-metr\" rel=\"nofollow noopener\" target=\"_blank\">Measured vs. forecast progress<\/a> \u00b7 <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7from-more-tokens-to-more-experiments\" rel=\"nofollow noopener\" target=\"_blank\">Diminishing returns<\/a><\/p>\n<p>People use \u201crecursive self-improvement\u201d to mean everything from AI boosting the productivity of human researchers to AI bootstrapping itself to incomprehensible intelligence. Here\u2019s my taxonomy: productivity gains (Type 1), increasing autonomy while still facing diminishing returns (Types 2\u20134), and a runaway loop to superintelligence if we can ever find accelerating returns (Type 5).<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!sTJ5!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5415417c-e2e1-4ed3-8ab7-c7a43552250d_1672x941.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/5415417c-e2e1-4ed3-8ab7-c7a43552250d_1672.jpeg\" width=\"1456\" height=\"819\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/5415417c-e2e1-4ed3-8ab7-c7a43552250d_1672x941.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:794056,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5415417c-e2e1-4ed3-8ab7-c7a43552250d_1672x941.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 2. Five types of AI self-improvement.<\/p>\n<p>We\u2019ve made real progress on Types 1 and 2: AI helps both researchers and engineers inside of AI companies, and powerful models can train and improve smaller ones. We haven\u2019t yet seen clear evidence for Type 3 (though <a href=\"https:\/\/www.alibabacloud.com\/en\/press-room\/alibaba-unveils-roadmap-on-full-stack-ai-strategy\" rel=\"nofollow noopener\" target=\"_blank\">Alibaba just made some strong claims<\/a>) and certainly not for Type 4. I do expect autonomous self-improvement to arrive at some point. I\u2019m skeptical that it leads to Type 5 &#8211; runaway super-intelligence &#8211; without a major conceptual breakthrough.<\/p>\n<p>There are plenty of other definitions of RSI, which can be a bit confusing. <a href=\"https:\/\/www.weco.ai\/blog\/4-levels-of-recursive-self-improvement\" rel=\"nofollow noopener\" target=\"_blank\">Weco\u2019s four levels of RSI<\/a> are close to mine. For a broader tour of all the things people mean when they say \u2018RSI\u2019, read <a href=\"https:\/\/tecunningham.github.io\/posts\/2026-06-05-rsi-definitions.html\" rel=\"nofollow noopener\" target=\"_blank\">Tom Cunningham\u2019s comprehensive guide<\/a>.<\/p>\n<p>I do expect narrow superintelligence in highly verifiable domains. Think chess, Go, formal math, parts of computer science and coding. Highly verifiable domains are largely formal and structured types of work where machines can generate unlimited training data, with perfect or near-perfect verification of correct vs incorrect, and do so entirely in software without waiting on the physical world or humans. That\u2019s an ideal setting for AI learning.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!ggzm!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd6ce012-aa5f-44a4-97eb-e3f30ae6d256_1672x941.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/cd6ce012-aa5f-44a4-97eb-e3f30ae6d256_1672.jpeg\" width=\"1456\" height=\"819\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/cd6ce012-aa5f-44a4-97eb-e3f30ae6d256_1672x941.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:550240,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd6ce012-aa5f-44a4-97eb-e3f30ae6d256_1672x941.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 3. What makes a domain highly verifiable?<\/p>\n<p>In fact, we already have narrow superintelligence in game plang. We\u2019re seeing it happen now in the most formal parts of math, in particular in proofs and in finding counter-examples that disprove major conjectures. For example, OpenAI recently reported <a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">an AI-generated proof resolving the Navier\u2013Stokes existence and smoothness problem<\/a>. Parts of software development are also extremely verifiable, while others are a bit less crisp (such as understanding what humans want).<\/p>\n<p>That isn\u2019t the same as broad superintelligence. Even our most powerful models need far more training data than humans, struggle to learn reliably from ongoing experience, and fail in surprising ways on tasks people find straightforward. Superhuman math doesn\u2019t automatically mean superhuman judgment everywhere else.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>Benchmarks and forecasts suggest that AI models should reliably succeed at coding tasks that take humans hours, without human help. The real world is messier. OpenAI\u2019s internal data shows much shorter stretches of autonomous work on research tasks.<\/p>\n<p>In its <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Research Acceleration \/ RSI report<\/a>, OpenAI showed how often its models completed tasks with and without human help, grouped by how long a human would need to do the work.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!vRrx!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fb54d7-f5cc-4471-990d-b32a54127900_3600x2770.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/29fb54d7-f5cc-4471-990d-b32a54127900_3600.jpeg\" width=\"1456\" height=\"1120\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/29fb54d7-f5cc-4471-990d-b32a54127900_3600x2770.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1120,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1690254,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29fb54d7-f5cc-4471-990d-b32a54127900_3600x2770.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 4. OpenAI\u2019s internal research tasks. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Even on tasks that would take a human less than 15 minutes, OpenAI\u2019s models succeeded without human intervention only 86% of the time. The estimated task length at 80% success was roughly 15 minutes over the first seven months of the year. July\u2019s results were similar to the whole period average.<\/p>\n<p>Fully autonomous RSI would require an AI to string together a great many research tasks reliably, stretching out over complex tasks that humans need weeks or months to accomplish. OpenAI\u2019s data suggests that we aren\u2019t close.<\/p>\n<p><a href=\"https:\/\/www.anthropic.com\/institute\/measuring-pace-of-ai-development\" rel=\"nofollow noopener\" target=\"_blank\">Anthropic also released a graph<\/a> showing how Claude accelerates AI research. It shows that internal AI models collaborate on or even lead more than 90% of R&amp;D tasks. That\u2019s objectively impressive. At the same time, the graph reports zero cases of AI autonomously completing AI R&amp;D tasks.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!L9j-!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a94e24-99c2-47a8-9dfb-fb8284aa98e3_1536x1024.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/35a94e24-99c2-47a8-9dfb-fb8284aa98e3_1536.jpeg\" width=\"1456\" height=\"971\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/35a94e24-99c2-47a8-9dfb-fb8284aa98e3_1536x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:371307,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35a94e24-99c2-47a8-9dfb-fb8284aa98e3_1536x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 5. Claude\u2019s role in internal AI R&amp;D. <a href=\"https:\/\/www.anthropic.com\/institute\/measuring-pace-of-ai-development\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>These are incredible tools. But they still need skilled people to set direction and get them back on track.<\/p>\n<p>For years, <a href=\"https:\/\/metr.org\/time-horizons\/\" rel=\"nofollow noopener\" target=\"_blank\">METR<\/a> has been publishing a chart showing what length of coding task (measured in human hours to complete) best-in-class AI models can achieve. It\u2019s been called <a href=\"https:\/\/benjamintodd.substack.com\/p\/the-most-important-graph-in-ai-right\" rel=\"nofollow noopener\" target=\"_blank\">the most important graph in AI<\/a>. METR\u2019s Mythos Preview evaluation estimated that the model could succeed at 80% of coding tasks that took humans three hours.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!IwIB!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3543dd81-587f-4697-9ff6-d48251fd4ce1_1448x1086.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3543dd81-587f-4697-9ff6-d48251fd4ce1_1448.jpeg\" width=\"1448\" height=\"1086\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3543dd81-587f-4697-9ff6-d48251fd4ce1_1448x1086.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1086,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:324836,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3543dd81-587f-4697-9ff6-d48251fd4ce1_1448x1086.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 6. METR\u2019s 80% task horizons. <a href=\"https:\/\/metr.org\/time-horizons\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p><a href=\"https:\/\/epoch.ai\/eci\" rel=\"nofollow noopener\" target=\"_blank\">Epoch\u2019s own rule of thumb<\/a> is that every five additional points of ECI (their overall benchmark of AI capability) correspond to roughly a doubling of METR\u2019s task horizon. Using that formula, we\u2019d expect GPT 5.6 Sol and GPT 6 Astra to be 80% successful at completing tasks of around 4 hours and 11 hours of human length, respectively.<\/p>\n<p>Another estimate (a forecast) of AI task length comes from the AI 2027 scenario, which estimated that by July 2026, frontier AIs would be 80% successful accomplishing tasks of around 11 hours. Fairly similar.<\/p>\n<p>The <a href=\"https:\/\/ai2027tracker.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2027 Tracker<\/a> charts all of these.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!orAt!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28331f2a-53ce-4cc9-b1f5-09aa2468b5d5_3600x1796.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/28331f2a-53ce-4cc9-b1f5-09aa2468b5d5_3600.jpeg\" width=\"1456\" height=\"726\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/28331f2a-53ce-4cc9-b1f5-09aa2468b5d5_3600x1796.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:726,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:574814,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28331f2a-53ce-4cc9-b1f5-09aa2468b5d5_3600x1796.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 7. The AI 2027 Tracker. <a href=\"https:\/\/ai2027tracker.com\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Inside OpenAI, though, the July research-task horizon at 80% success was roughly 15 minutes.<\/p>\n<p>Here\u2019s the gap:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!LKkm!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6e52468-d2e1-4e9a-8349-eba4e15862f4_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f6e52468-d2e1-4e9a-8349-eba4e15862f4_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f6e52468-d2e1-4e9a-8349-eba4e15862f4_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:854528,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6e52468-d2e1-4e9a-8349-eba4e15862f4_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 8. Forecasts, benchmarks, and real AI research. <a href=\"https:\/\/ai2027tracker.com\/\" rel=\"nofollow noopener\" target=\"_blank\">Tracker<\/a> \u00b7 <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI<\/a>.<\/p>\n<p>A four-hour benchmark horizon is about 16 times longer than OpenAI\u2019s research horizon. AI 2027\u2019s 11-hour forecast is about 44 times longer. Of course, the tasks being performed by researchers at OpenAI aren\u2019t the same as those in the METR benchmark. So we should expect some discrepancy. This, however, goes well beyond that.<\/p>\n<p>Actual AI research at OpenAI is an order of magnitude or more harder than metrics, benchmarks, or forecasts suggest. That should make us wary of relying too much on benchmarks, or of saying that future scenarios like AI 2027 are \u2018on track.\u2019 The authors of the related <a href=\"https:\/\/ai-2040.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2040 project<\/a> still describe AI 2027 as roughly the future they expect, and say reality is tracking closer to it than even they expected. That\u2019s not what we see from within OpenAI. This isn\u2019t an apples-to-apples comparison, but the difference is remarkable. AI 2027 appears to be substantially over-optimistic in this regard.<\/p>\n<p>In January of this year, <a href=\"https:\/\/arachnemag.substack.com\/p\/the-metr-graph-is-hot-garbage\" rel=\"nofollow noopener\" target=\"_blank\">Nathan Witkin made a case that the METR graph was exaggerating progress<\/a>. The real world data suggests that at least some of his critiques were correct. The gap between benchmarks, forecasts, and data gleaned from actual use of AI should influence our expectations about the future.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>OpenAI\u2019s report also shows impressive increases in AI token usage, in compute spend per researcher, and in lines of code written. But these aren\u2019t results. They\u2019re intermediate measures. How much progress do they actually drive?<\/p>\n<p>Researchers used 124x more tokens per person. Engineers shipped roughly 7x as many lines of code per person. Researchers ran 1.6x as many experiments per researcher vs OpenAI\u2019s 2025 whole year average.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!0k7X!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e5df740-835c-4a8d-987b-d8561ec0172d_3600x2109.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7e5df740-835c-4a8d-987b-d8561ec0172d_3600.jpeg\" width=\"1456\" height=\"853\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7e5df740-835c-4a8d-987b-d8561ec0172d_3600x2109.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:853,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1075582,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e5df740-835c-4a8d-987b-d8561ec0172d_3600x2109.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 9. Token use inside OpenAI. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!wXPv!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d63908d-9276-4f24-9f28-fa26a5c855f9_3600x2138.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/6d63908d-9276-4f24-9f28-fa26a5c855f9_3600.jpeg\" width=\"1456\" height=\"865\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/6d63908d-9276-4f24-9f28-fa26a5c855f9_3600x2138.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:865,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1061007,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d63908d-9276-4f24-9f28-fa26a5c855f9_3600x2138.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 10. Experiment pace inside OpenAI. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!Br45!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d43599-d804-4c46-87bb-c9ffe301a1f3_3600x762.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/74d43599-d804-4c46-87bb-c9ffe301a1f3_3600.jpeg\" width=\"1456\" height=\"308\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/74d43599-d804-4c46-87bb-c9ffe301a1f3_3600x762.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:308,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:327950,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d43599-d804-4c46-87bb-c9ffe301a1f3_3600x762.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 11. From tokens to code to experiments. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>More tokens and code don\u2019t tell us much on their own. The 1.6\u00d7 experiment pace is closer to useful research output. Even that doesn\u2019t mean AI is improving 1.6\u00d7 faster.<\/p>\n<p>An enormous increase in AI output has accompanied a much smaller increase in experiments run.<\/p>\n<p>This isn\u2019t a controlled experiment. We don\u2019t know what would happen if researchers switched back to an older model. But it gives us a useful view of AI-assisted research inside a frontier lab.<\/p>\n<p>It\u2019s not just OpenAI. Anthropic reports that their engineers are now <a href=\"https:\/\/www.anthropic.com\/institute\/recursive-self-improvement\" rel=\"nofollow noopener\" target=\"_blank\">producing 8x as many lines of code<\/a> per person as they did in 2024 &#8211; somewhat similar to OpenAI. Anthropic also sees significant diminishing returns between productivity and AI progress. Here\u2019s a direct quote from its <a href=\"https:\/\/www-cdn.anthropic.com\/7624816413e9b4d2e3ba620c5a5e091b98b190a5\/Claude%20Mythos%20Preview%20System%20Card.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Mythos Preview system card<\/a>:<\/p>\n<p>\u201cProductivity uplift does not translate one-for-one to capabilities progress. We surveyed technical staff on the productivity uplift they experience from Claude Mythos Preview relative to zero AI assistance. The distribution is wide and the geometric mean is on the order of 4\u00d7. [\u2026] We estimate that reaching 2\u00d7 on overall progress via this channel would require uplift roughly an order of magnitude larger than what we observe.\u201d- Anthropic, Claude Mythos Preview System Card; emphasis mine<\/p>\n<p>Translation: To double the pace of AI progress, Anthropic estimates that AI would need to increase the productivity of their employees by roughly a factor of 40 relative to no AI assistance.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!9v0h!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7fc24dad-736a-4bf6-8ddf-c23edf15ebd2_2700x1050.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7fc24dad-736a-4bf6-8ddf-c23edf15ebd2_2700.jpeg\" width=\"1456\" height=\"566\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/7fc24dad-736a-4bf6-8ddf-c23edf15ebd2_2700x1050.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:566,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:465913,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7fc24dad-736a-4bf6-8ddf-c23edf15ebd2_2700x1050.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 12. Anthropic\u2019s productivity-to-progress estimate. <a href=\"https:\/\/www.anthropic.com\/institute\/recursive-self-improvement\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>This is an estimate, not a measurement of progress. Even the 4\u00d7 productivity figure comes from an <a href=\"https:\/\/www.anthropic.com\/institute\/recursive-self-improvement\" rel=\"nofollow noopener\" target=\"_blank\">opt-in survey of 130 Anthropic staff<\/a>. I put more weight on OpenAI\u2019s logged experiments, though the two sources measure different things.<\/p>\n<p>We don\u2019t yet know how much those extra experiments are accelerating AI improvement, if at all. In general, there are also <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A77-why-does-progress-get-harder\" rel=\"nofollow noopener\" target=\"_blank\">steeply diminishing returns of more experiments<\/a> in most branches of science. That means that a 60% increase in experiment pace could be on the order of a 10% boost to AI improvement pace. (A power law exponent of 0.2, for those who want to do the math.) That\u2019s speculation for now. We\u2019ll learn more as the labs publish results.<\/p>\n<p>What about giving the same AI model more time to think?<\/p>\n<p>That scales badly also. In <a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI\u2019s recently publicized results on unsolved math problems<\/a>, success rises roughly with the log of compute over the range shown. It shows logarithmic diminishing returns. In plain English, each additional doubling of compute for a model buys roughly the same gain in success rate, while costing twice as much.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!Axtn!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa744791b-ec7d-4022-b83a-464b51fa5eec_3600x2640.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a744791b-ec7d-4022-b83a-464b51fa5eec_3600.jpeg\" width=\"1456\" height=\"1068\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a744791b-ec7d-4022-b83a-464b51fa5eec_3600x2640.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1068,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:763826,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa744791b-ec7d-4022-b83a-464b51fa5eec_3600x2640.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 13. Test-time compute and math performance. <a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>What if we throw more agents at it instead? A common RSI \/ ASI idea is that once we have AIs at a certain capability level, we can just spawn more copies and put them to work.<\/p>\n<p>Adding agents can get tasks done faster and sometimes reach a higher capability level. But on the three benchmarks in <a href=\"https:\/\/www.tobyord.com\/writing\/swarm-scaling\" rel=\"nofollow noopener\" target=\"_blank\">Toby Ord\u2019s analysis<\/a>, expanding a swarm buys less improvement per token than letting one agent think longer.<\/p>\n<p>His rough rule of thumb is a square root. If one agent can accomplish a task in 10 hours, then 100 agents could accomplish it in one hour. The speedup is 10, the square root of the number of agents (100). But to get this speedup, you increase the total cost in tokens or run time compute by the same factor. So going from one to 100 agents can get a task done in one tenth the time. But it\u2019ll be ten times as expensive.<\/p>\n<p>Parallel agents can save time, at a much higher compute cost.<\/p>\n<p>Another challenge is that agents often think alike. In a <a href=\"https:\/\/arxiv.org\/abs\/2412.03151\" rel=\"nofollow noopener\" target=\"_blank\">study comparing LLMs with 467 people<\/a>, the first ten AI responses offered collective creativity comparable to about eight to ten people. After that, roughly two extra AI responses added as much as one extra human response. A <a href=\"https:\/\/arxiv.org\/abs\/2501.19361\" rel=\"nofollow noopener\" target=\"_blank\">separate study across model families<\/a> also found less diversity in AI responses. That doesn\u2019t mean every agent has the same idea. But a hundred copies may offer less variety than a hundred different researchers.<\/p>\n<p>None of this makes swarms useless-or safe. <a href=\"https:\/\/scaling01.substack.com\/p\/accidental-scaling\" rel=\"nofollow noopener\" target=\"_blank\">Lisan al-Gaib makes a strong case for parallel agent swarms as a potent cyber-weapon in \u201cAccidental Scaling.\u201d<\/a> I don\u2019t share all of his assessment of what swarms have accomplished. In math, for example, I think he gives far too much credit to the swarm and not enough to the better internal model that OpenAI used.<\/p>\n<p><a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI says the model behind its Navier\u2013Stokes result<\/a> was developed through \u201clarge-scale reinforcement learning on top of a previously pretrained model.\u201d <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7we-already-have-narrow-superintelligence\" rel=\"nofollow noopener\" target=\"_blank\">Formal math is a highly verifiable domain<\/a>, which makes it a particularly good fit for that approach: Machines can generate nearly limitless amounts of training data, and verify that solutions are correct or incorrect, all in software. My guess is that this model\u2019s full results will show an especially large improvement in math.<\/p>\n<p><a href=\"https:\/\/www.dwarkesh.com\/p\/noam-brown\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI\u2019s Noam Brown made the central point explicitly<\/a>: he wouldn\u2019t give multi-agent methods even 10% of the credit for the Navier\u2013Stokes result.<\/p>\n<p>I do think Lisan makes good points about cybersecurity. If you\u2019re searching for a security vulnerability at a target site and can divide the search among agents, speed may justify a huge token bill. Swarms can be dangerous even when they\u2019re inefficient.<\/p>\n<p>I\u2019m less convinced that this scales to research breakthroughs. Inventing something like the transformer probably takes more than searching a space someone has already defined.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>Building a better model can bring gains that extra thinking time or more copies of the old model can\u2019t. Look at the gap between Astra and OpenAI\u2019s internal model on the same math problems.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!wN_k!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4dd75a51-5b82-4cb7-84cc-f917deceeaf6_3600x2138.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/4dd75a51-5b82-4cb7-84cc-f917deceeaf6_3600.jpeg\" width=\"1456\" height=\"865\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/4dd75a51-5b82-4cb7-84cc-f917deceeaf6_3600x2138.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:865,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1005703,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4dd75a51-5b82-4cb7-84cc-f917deceeaf6_3600x2138.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 14. Better models versus more thinking time. <a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>That\u2019s the strongest version of the RSI argument: a more capable AI could do research that today\u2019s model can\u2019t do, however many copies we run.<\/p>\n<p>But building that better model also runs into diminishing returns. More training data, more training compute, larger models, and more reinforcement-learning (RL) compute all show diminishing returns in published scaling studies. Making dense models larger usually raises the compute needed for each output token, too. None of these routes gives us a free pass around the problem.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!hxKh!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb192c5bf-9354-47b8-8477-42ebd568f424_3600x1580.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/b192c5bf-9354-47b8-8477-42ebd568f424_3600.jpeg\" width=\"1456\" height=\"639\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/b192c5bf-9354-47b8-8477-42ebd568f424_3600x1580.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:639,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:970037,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb192c5bf-9354-47b8-8477-42ebd568f424_3600x1580.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 15. Diminishing returns to scaling. <a href=\"https:\/\/arxiv.org\/abs\/2203.15556\" rel=\"nofollow noopener\" target=\"_blank\">Chinchilla<\/a> \u00b7 <a href=\"https:\/\/openreview.net\/pdf\/73a75bec6ecd647f8a93556e89bd4ed972c80ce1.pdf\" rel=\"nofollow noopener\" target=\"_blank\">ScaleRL<\/a> \u00b7 <a href=\"https:\/\/openai.com\/index\/navier-stokes-solution\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI<\/a>.<\/p>\n<p>Those scaling results give us reason to expect diminishing returns when AI helps build the next model, too.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>AI capabilities are rising quickly. But the public data doesn\u2019t show a sustained acceleration. To the extent that AI tools are boosting productivity, they may be being offset by <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A77-why-does-progress-get-harder\" rel=\"nofollow noopener\" target=\"_blank\">the problems growing harder<\/a>. Or we may simply be early. Either way, the trend isn\u2019t showing a fast takeoff.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!26CF!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe921291b-29e4-46eb-8565-7f938566fa01_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/e921291b-29e4-46eb-8565-7f938566fa01_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/e921291b-29e4-46eb-8565-7f938566fa01_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:827177,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe921291b-29e4-46eb-8565-7f938566fa01_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 16. Frontier ECI gains since January 2024. <a href=\"https:\/\/epoch.ai\/eci\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>The <a href=\"https:\/\/epoch.ai\/eci\" rel=\"nofollow noopener\" target=\"_blank\">public ECI frontier<\/a>-the best score among models released by each date-has gained about 16 points a year on a trend fitted from January 2024 through September 2026. That\u2019s blisteringly fast progress, but this period doesn\u2019t show a runaway surge.<\/p>\n<p>Here\u2019s the same frontier in absolute ECI points, through July 2026, to put it in perspective.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!dSEr!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaf88f4-62d2-4bff-b0ca-1e6b1a20aebf_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3eaf88f4-62d2-4bff-b0ca-1e6b1a20aebf_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3eaf88f4-62d2-4bff-b0ca-1e6b1a20aebf_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:762472,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eaf88f4-62d2-4bff-b0ca-1e6b1a20aebf_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 17. The absolute frontier ECI score. <a href=\"https:\/\/epoch.ai\/eci\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>The public frontier also can\u2019t tell us everything happening inside the labs. Anthropic gives us a closer look in the <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Opus 5.5 system card<\/a>, using its own version of the index, AECI.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!TfIz!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feadf1eaf-d39b-4504-ad22-f8e41f9e0041_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/eadf1eaf-d39b-4504-ad22-f8e41f9e0041_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/eadf1eaf-d39b-4504-ad22-f8e41f9e0041_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:941463,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feadf1eaf-d39b-4504-ad22-f8e41f9e0041_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 18. Anthropic\u2019s fitted capability trend. <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p><a href=\"https:\/\/aifuturesnotes.substack.com\/p\/on-mythoss-ai-r-and-d-abilities\" rel=\"nofollow noopener\" target=\"_blank\">Eli Lifland<\/a>, a co-author of <a href=\"https:\/\/ai-2027.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2027<\/a> and <a href=\"https:\/\/ai-2040.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2040<\/a>, saw the apparent trend break as a warning that we were heading toward an intelligence explosion:<\/p>\n<p>\u201cAnthropic is probably right here [that they hadn\u2019t reached dangerous levels of AI self-improvement], but alarm bells should be going off! Our processes are not ready to handle an intelligence explosion and we appear to be going full-steam ahead toward one.\u201d<br \/>&#8211; Eli Lifland, <a href=\"https:\/\/aifuturesnotes.substack.com\/p\/on-mythoss-ai-r-and-d-abilities\" rel=\"nofollow noopener\" target=\"_blank\">On Mythos\u2019s AI R&amp;D Capabilities<\/a><\/p>\n<p>What looked like acceleration now appears more consistent with a one-time jump. The level went up. The rate hasn\u2019t kept climbing.<\/p>\n<p>Achieving those gains has required an enormous increase in the inputs to AI. For example, consider computing power. <a href=\"https:\/\/epoch.ai\/data-insights\/ai-chip-production\" rel=\"nofollow noopener\" target=\"_blank\">Epoch\u2019s estimates of AI chip capacity<\/a>, measured in NVIDIA H100 equivalents, show roughly 127-fold growth in just over three years (including projections at the end of this period).<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!7jG_!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70b1e1b5-ad21-4ea5-ab5d-8ac90db599cb_3600x2026.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/70b1e1b5-ad21-4ea5-ab5d-8ac90db599cb_3600.jpeg\" width=\"1456\" height=\"819\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/70b1e1b5-ad21-4ea5-ab5d-8ac90db599cb_3600x2026.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:792117,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70b1e1b5-ad21-4ea5-ab5d-8ac90db599cb_3600x2026.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 19. AI chip capacity and frontier ECI. <a href=\"https:\/\/epoch.ai\/data-insights\/ai-chip-production\" rel=\"nofollow noopener\" target=\"_blank\">Source: Epoch AI<\/a>.<\/p>\n<p>This is total AI chip capacity, including inference. Still, the increase is striking: vastly more computing capacity has accompanied much steadier gains in measured capability.<\/p>\n<p>The broader picture looks similar. Here are six inputs alongside capability gains, going back to February 2023.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!mRwm!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0e10238-ffa7-4285-a550-e918adc4c748_3600x3114.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/c0e10238-ffa7-4285-a550-e918adc4c748_3600.jpeg\" width=\"1456\" height=\"1259\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/c0e10238-ffa7-4285-a550-e918adc4c748_3600x3114.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1259,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1730526,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc0e10238-ffa7-4285-a550-e918adc4c748_3600x3114.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 20. Six inputs alongside frontier ECI. <a href=\"https:\/\/epoch.ai\/data-insights\/ai-chip-production\" rel=\"nofollow noopener\" target=\"_blank\">Epoch chip data<\/a> \u00b7 <a href=\"https:\/\/newsletter.semianalysis.com\/p\/long-live-the-short-king-why-4-hi\" rel=\"nofollow noopener\" target=\"_blank\">SemiAnalysis workload shares<\/a>.<\/p>\n<p>Everywhere we look, AI has diminishing returns. It gets more expensive in treasure and talent to make each step forward. More of every input has been required to maintain steady gains in AI capabilities.<\/p>\n<p>We\u2019ve been able to scale these inputs because, until recently, the cost was within the scope of what hyperscalers could pay from their profits. That is no longer the case. From this point forward, future AI investment will increasingly depend on AI revenues going up. And the scale of the numbers &#8211; 3% of US GDP is now going into AI infrastructure &#8211; suggests that eventually the growth rate will decline. If investment growth does slow, to anything less than its current blistering exponential pace, capability progress could slow too. Even if investment growth continues (which I expect for the foreseeable future) a slowdown from its current exponential growth rate to a more modest one (which I also expect) could lead to a slower pace of progress. Better AI research tools may be needed to offset that.<\/p>\n<p>The day when we need better AI tools just to continue the pace of AI progress may already have arrived. Not because investment is slowing, but because the problem of improving AI itself gets harder at each step.<\/p>\n<p>Here\u2019s Anthropic in the <a href=\"https:\/\/www.anthropic.com\/claude-fable-5-1-mythos-5-1-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Mythos 5.1 system card<\/a>:<\/p>\n<p>\u201cwe believe that internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate.\u201d- Anthropic, Claude Fable 5.1 &amp; Claude Mythos 5.1 System Card, section 2.3 \u2013 emphasis theirs.<\/p>\n<p>The key word is maintaining-and Anthropic italicized that word in its own system card. Increasingly capable AI may be essential just to keep the pace of improvement where it is.<\/p>\n<p>Opus 5.5 improves substantially on several coding and computer use benchmarks. But on CoBench, Anthropic\u2019s benchmark built from historical AI R&amp;D problems, it gains just 2.6 percentage points over Opus 5, within the reported error bars.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!McvY!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95f8d492-14ec-4bd1-a606-d333f82afd6a_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/95f8d492-14ec-4bd1-a606-d333f82afd6a_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/95f8d492-14ec-4bd1-a606-d333f82afd6a_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:905659,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95f8d492-14ec-4bd1-a606-d333f82afd6a_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 21. Opus 5.5 benchmark gains. <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Why the smaller gain here? Maybe AI research is simply harder than other tasks. Bear in mind that CoBench isn\u2019t testing the ability to produce significant discoveries. It\u2019s much more limited in scope. It asks models to investigate historical AI R&amp;D problems using code, logs, and documents. That\u2019s useful research debugging and productivity work, but it doesn\u2019t directly test whether a model can invent a new architecture or make a conceptual breakthrough.<\/p>\n<p>The evidence on open-ended research suggests another obstacle: coming up with useful ideas that haven\u2019t already been tried.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>Why do useful new ideas often get harder to find?<\/p>\n<p>Tom Cunningham and Manish Shetty have a useful <a href=\"https:\/\/tecunningham.github.io\/posts\/2026-03-13-apple-picking-ai.html\" rel=\"nofollow noopener\" target=\"_blank\">apple-picking metaphor<\/a>. An AI can pick the low-hanging fruit quickly, while humans can still reach ideas the AI can\u2019t.<\/p>\n<p>Once those apples are picked, another copy of the same agent finding them again doesn\u2019t help. A stronger model can reach higher. To add my own flourish, the apples may also get sparser and farther apart as you climb. The RSI question is whether each harvest gives us enough to build a better apple-picker.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!6bOD!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6551a7ec-6ae3-4b0e-81cf-2401cf240736_3600x1810.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/6551a7ec-6ae3-4b0e-81cf-2401cf240736_3600.jpeg\" width=\"1456\" height=\"732\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/6551a7ec-6ae3-4b0e-81cf-2401cf240736_3600x1810.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:732,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:459413,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6551a7ec-6ae3-4b0e-81cf-2401cf240736_3600x1810.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 22. The apple-picking model of AI R&amp;D. <a href=\"https:\/\/tecunningham.github.io\/posts\/2026-03-13-apple-picking-ai.html\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>This pattern shows up across R&amp;D. <a href=\"https:\/\/www.aeaweb.org\/articles?id=10.1257\/aer.20180338\" rel=\"nofollow noopener\" target=\"_blank\">Bloom and colleagues<\/a> document fields where research effort grows while research productivity falls. A famous example is Eroom\u2019s Law: in the historical drug-development data, the inflation-adjusted R&amp;D cost per new approved drug roughly doubled every nine years.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!fX2g!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68197d17-1bf6-4a72-b06c-db60987eb10a_3600x2410.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/68197d17-1bf6-4a72-b06c-db60987eb10a_3600.jpeg\" width=\"1456\" height=\"975\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/68197d17-1bf6-4a72-b06c-db60987eb10a_3600x2410.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:975,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1099154,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68197d17-1bf6-4a72-b06c-db60987eb10a_3600x2410.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 23. Eroom\u2019s Law in drug development. <a href=\"https:\/\/www.nature.com\/articles\/nrd3681\/figures\/1\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Pharma has other complications, including regulation, difficult clinical trials, and rising expectations for safety. Existing treatments can also raise the bar for a useful new drug. But some of this difficulty may also be that the low-hanging fruit has been picked.<\/p>\n<p>Stockfish, the chess engine, gives us a more direct look at software research. We have records of experiments aimed at improving it and the gains that followed. This gives us a real-world dataset to look at the gains of experimentation in software. As a result, several RSI models draw on this data. That said, not all the improvements came from these experiments. Several important ideas also came from outside the project, so we shouldn\u2019t give its experiments all the credit.<\/p>\n<p><a href=\"https:\/\/epoch.ai\/publications\/do-the-returns-to-software-rnd-point-towards-a-singularity\" rel=\"nofollow noopener\" target=\"_blank\">Epoch\u2019s analysis of software R&amp;D<\/a> estimates returns to research effort at about 0.83 for Stockfish, a bit slower than linear. These are diminishing returns, but gentle ones. These returns, however, are improvements in computational efficiency. And more compute does not turn directly into more AI capability. As we saw earlier, <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A75-better-models-matter-more-than-more-copies\" rel=\"nofollow noopener\" target=\"_blank\">AI capability also has steep diminishing returns from adding more computational power<\/a>. So we shouldn\u2019t read that 0.83 as the return from experimentation to AI capability itself. AI capability grows much more slowly than compute, as we\u2019ve seen already.<\/p>\n<p>Andrej Karpathy\u2019s <a href=\"https:\/\/github.com\/karpathy\/autoresearch\" rel=\"nofollow noopener\" target=\"_blank\">autoresearch demonstration<\/a> gets closer to the process we want to understand. A \u201cteacher\u201d AI agent changes a smaller \u201cstudent\u201d AI model\u2019s training code, runs it, checks the result, and tries again. The teacher agent itself doesn\u2019t improve, but it is able to improve the \u201clearner\u201d. This is my <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A72-what-does-rsi-mean\" rel=\"nofollow noopener\" target=\"_blank\">Type 2: A stronger AI improves a weaker one<\/a>.<\/p>\n<p><a href=\"https:\/\/github.com\/karpathy\/autoresearch\/discussions\/32\" rel=\"nofollow noopener\" target=\"_blank\">One public run<\/a>, posted by an agent operating on Karpathy\u2019s behalf, reported 89 experiments over roughly 7.5 hours. About 92% of that session\u2019s gain arrived by run 44. Gains came quickly, then slowed. The setup was deliberately small, with a five-minute training budget per experiment. But the agent could change the architecture, optimizer, and training settings; it wasn\u2019t limited to a handful of knobs.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!-U_r!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ff71dd9-31d4-4314-a821-506435af1dc5_3600x2026.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/4ff71dd9-31d4-4314-a821-506435af1dc5_3600.jpeg\" width=\"1456\" height=\"819\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/4ff71dd9-31d4-4314-a821-506435af1dc5_3600x2026.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:839152,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ff71dd9-31d4-4314-a821-506435af1dc5_3600x2026.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 24. Gains in one autoresearch run. <a href=\"https:\/\/github.com\/karpathy\/autoresearch\/discussions\/32\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>A<a href=\"https:\/\/github.com\/karpathy\/autoresearch\/discussions\/43\" rel=\"nofollow noopener\" target=\"_blank\"> later public run<\/a> got further, so the first run hadn\u2019t hit a hard ceiling. This is a useful early example of autonomous research, and yet another place where we see the diminishing returns endemic in AI research. That said, this was a very early experiment. I expect future systems to do much better. This particular AI improvement loop will likely grow stronger.<\/p>\n<p>This is where the distinction matters. More tokens can buy more code, and more code can help us run more experiments. But experiments only improve AI if they uncover something useful.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!8ZPD!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf49dc6c-cc25-48aa-9d9b-d0969a4a2457_3600x2760.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/df49dc6c-cc25-48aa-9d9b-d0969a4a2457_3600.jpeg\" width=\"1456\" height=\"1116\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/df49dc6c-cc25-48aa-9d9b-d0969a4a2457_3600x2760.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1116,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1869350,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf49dc6c-cc25-48aa-9d9b-d0969a4a2457_3600x2760.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 25. From AI activity to useful improvements.<\/p>\n<p>The bigger question is whether AI can come up with ambitious new research ideas or conceptual breakthroughs.<\/p>\n<p>Anthropic\u2019s description of Opus 5.5 is blunt:<\/p>\n<p>\u201cAs with previous models, it is weaker on open-ended research: internal users report that it mostly tests incremental ideas and prefers less ambitious hypotheses, and in our human-run biology exercise, it deferred to the published literature and struggled to develop novel ideas (Section 2.2.2).\u201d- Anthropic, <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Claude Opus 5.5 System Card<\/a>, section 2.3.3; emphasis mine<\/p>\n<p>METR\u2019s assessment in the same card identifies what may still be missing:<\/p>\n<p>\u201cThis is highly uncertain, but we expect that full automation of AI R&amp;D will require large improvements in foresight, prediction, creating one\u2019s own feedback loops, and generally other skills that might typically be referred to as researcher \u2018judgement\u2019 or \u2018taste\u2019.\u201d- METR, quoted in the <a href=\"https:\/\/www.anthropic.com\/claude-opus-5-5-system-card\" rel=\"nofollow noopener\" target=\"_blank\">Claude Opus 5.5 System Card<\/a>, section 2.3.6<\/p>\n<p>In these examples, humans still supply much of the direction and judgment.<\/p>\n<p>Future models will probably get better at this. But in the world\u2019s stockpile of potential training data, we have many more examples of incremental work than of breakthroughs. I wonder whether that makes novelty harder to learn. That\u2019s speculation, but worth watching.<\/p>\n<p>This is also tough to address by simply running more copies of the AI. A huge number of parallel agents can help with the incremental improvements or searching over a large set of parameters, but for breakthrough ideas they may run into the homogeneity problem: <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7what-about-agent-swarms\" rel=\"nofollow noopener\" target=\"_blank\">More parallel agents still think alike<\/a>.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>How far are we from the self-improvement loop being strong enough to sustain itself, or to propel itself into runaway super-intelligence? Can we quantify this?<\/p>\n<p>We can make a rough estimate. Better AI helps with research; useful research produces better AI. For the loop to sustain itself, each round must produce enough gains to propel the system through the next loop, even as improvements get harder to discover.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!djgS!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f3abc4f-496a-43a7-90f2-05955d2c6cff_3600x3220.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/8f3abc4f-496a-43a7-90f2-05955d2c6cff_3600.jpeg\" width=\"1456\" height=\"1302\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/8f3abc4f-496a-43a7-90f2-05955d2c6cff_3600x3220.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1302,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1747155,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f3abc4f-496a-43a7-90f2-05955d2c6cff_3600x3220.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 26. The AI self-improvement loop. <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Model<\/a>.<\/p>\n<p>In a recent paper, <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">The Economics of Recursive Self-Improvement<\/a>, Tom Cunningham and colleagues modeled this from the standpoint of how much more productivity every point of additional ECI produces from an AI. They ask first and foremost what that number would need to be to create a self-sustaining feedback loop. And secondly, they try to determine what that productivity-per-ECI-point number is today.<\/p>\n<p>First, they find a self-sustaining RSI threshold of roughly 15% more research productivity per extra ECI point. In their model, that\u2019s about where better AI would generate enough progress to sustain the loop.<\/p>\n<p>The picture below shows the idea. At the threshold, each cycle of gains powers the next. Above the threshold, the feedback loop accelerates. Below the threshold, the feedback loop is too weak, and the rate of improvement it brings drops on each cycle. This model isolates the software loop; outside investment can still drive rapid progress.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!s-s9!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1873d1fe-339c-47d8-bcce-12c0a8cfb39b_3600x2120.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/1873d1fe-339c-47d8-bcce-12c0a8cfb39b_3600.jpeg\" width=\"1456\" height=\"857\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/1873d1fe-339c-47d8-bcce-12c0a8cfb39b_3600x2120.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:857,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:907441,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1873d1fe-339c-47d8-bcce-12c0a8cfb39b_3600x2120.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 27. Three illustrative feedback paths. <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Updating this slightly with <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7lessons-from-software-r-and-d\" rel=\"nofollow noopener\" target=\"_blank\">data from the Stockfish experiments<\/a> puts the threshold a little higher, at roughly 19% per ECI point. I wouldn\u2019t put much weight on that precise difference. Both estimates are uncertain. But they give us a way to think about the strength of the feedback loop and a rough band at which self-sustaining or runaway RSI may begin.<\/p>\n<p>The second thing Cunningham and team do is make a rough estimate that the current AI productivity gain is about 9% per ECI point. That\u2019s below their self-sustaining threshold.<\/p>\n<p>I like the model. OpenAI\u2019s newer data, however, suggests the loop may be quite a bit weaker.<\/p>\n<p>Cunningham\u2019s estimate of 9% productivity gain per ECI point is based on Anthropic\u2019s survey of 130 staff, who reported roughly 4\u00d7 the productivity they\u2019d have without AI. Cunningham and colleagues compare that with a 16-point capability gain since early Claude Code.<\/p>\n<p>That comparison assumes the earlier tools added little or no productivity, so \u2018no AI\u2019 is a reasonable starting point. The authors say this explicitly. I\u2019m not sure the assumption holds for the same researchers doing the same work, but that\u2019s a smaller issue.<\/p>\n<p>The authors themselves know that this is a rough calculation, and warn that the 4\u00d7 survey estimate is probably too high.<\/p>\n<p>OpenAI\u2019s newer data gives us a firmer way to check the number: Actual logged experiments over time, rather than human estimates of their own productivity with and without AI. I put more weight on this for three reasons:<\/p>\n<p>Direct and broad measurement. Instead of relying on surveys, <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI actually tracked and measured experiments run on their infrastructure<\/a>. That means they didn\u2019t rely on researchers estimating their own productivity, which can be far off.<\/p>\n<p>Full sample, not opt-in. Similarly, OpenAI\u2019s data catches every active experimenter, while Anthropic\u2019s only reflects the 130 employees who took the time to answer the survey \u2013 and who therefore may not be a representative set.<\/p>\n<p>Enormously more data. We don\u2019t know how many experiments are in the 32 weeks of OpenAI data, but it\u2019s likely at least tens of thousands of individual examples and possibly hundreds of thousands.<\/p>\n<p>Any way you slice it, the new OpenAI data, released after Cunningham\u2019s paper was drafted, is a larger, more comprehensive, more representative, and almost certainly more accurate dataset than Anthropic\u2019s internal opt-in survey of employees.<\/p>\n<p>Now let\u2019s use OpenAI\u2019s experiment data to calibrate the productivity gain per ECI point. We know that in August, OpenAI researchers ran ~1.6\u00d7 as many experiments per person per month as the 2025 average. If we pair that with <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A76-were-not-seeing-runaway-acceleration\" rel=\"nofollow noopener\" target=\"_blank\">roughly 16 points of frontier ECI improvement<\/a>, it works backward to about 3% productivity gain per point of ECI. By contrast, 9% compounded over 16 points would mean roughly 4\u00d7 productivity.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!aY6l!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff575fdcf-98c1-415b-baa0-c0d7cf480504_3600x2100.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f575fdcf-98c1-415b-baa0-c0d7cf480504_3600.jpeg\" width=\"1456\" height=\"849\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f575fdcf-98c1-415b-baa0-c0d7cf480504_3600x2100.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:849,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1068044,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff575fdcf-98c1-415b-baa0-c0d7cf480504_3600x2100.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 28. Comparing productivity estimates. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI methods<\/a>.<\/p>\n<p>Here\u2019s OpenAI\u2019s published weekly series alongside that hypothetical path of 9% more productivity per additional ECI point. The blue line ends at ~1.6\u00d7. The red line shows what 9% per point would imply if 16 ECI points were spread across this period. That doesn\u2019t match what we see from OpenAI\u2019s data. I want to be clear here that all data sets are noisy. We don\u2019t know exactly what model researchers were using on what days, or whether the new experiments were also higher quality than old experiments. We need more experiments and more data to further calibrate these numbers. Working with what we do have, what we see is a quite low boost to productivity from each additional ECI point.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!b1_h!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3978656d-0208-4928-95c3-672974275431_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3978656d-0208-4928-95c3-672974275431_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/3978656d-0208-4928-95c3-672974275431_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:903094,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3978656d-0208-4928-95c3-672974275431_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 29. Experiment pace versus a hypothetical path. <a href=\"https:\/\/openai.com\/index\/research-acceleration-view-inside-openai\/\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>Even that 3% could give better models too much credit. <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7from-more-tokens-to-more-experiments\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI also used far more tokens<\/a> and had more compute for experiments. Those could account for some of the increase in experiment pace. So the range is probably a bit lower.<\/p>\n<p>I use 2\u20133% productivity gain per ECI point as a working assumption, allowing for some help from those other inputs. This is still a rough estimate, albeit one that\u2019s based on the best real-world data we have.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!trW9!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc8fe56-53e0-4628-8c37-de6fc03ad842_3600x2136.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/ccc8fe56-53e0-4628-8c37-de6fc03ad842_3600.jpeg\" width=\"1456\" height=\"864\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/ccc8fe56-53e0-4628-8c37-de6fc03ad842_3600x2136.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:924542,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc8fe56-53e0-4628-8c37-de6fc03ad842_3600x2136.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 30. Productivity estimates and the takeoff threshold. <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>With those assumptions, 2\u20133% per ECI point against a 15\u201319% threshold leaves a roughly five- to tenfold gap. That\u2019s a big gap, though its size depends on how well experiment counts capture useful research and whether the assumed capability change is right.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!A2Jg!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f3c2459-0127-4fdc-b762-49b6cd1671ce_1672x993.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/0f3c2459-0127-4fdc-b762-49b6cd1671ce_1672.jpeg\" width=\"1456\" height=\"865\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/0f3c2459-0127-4fdc-b762-49b6cd1671ce_1672x993.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:865,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:637496,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f3c2459-0127-4fdc-b762-49b6cd1671ce_1672x993.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 31. Diminishing returns around the loop. <a href=\"https:\/\/elasticity.institute\/rsi-paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>AI is helping build better AI. Under this estimate, though, each turn of the loop adds less than the last. The feedback would have to become much stronger to sustain itself.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>This software loop sits alongside <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7keeping-up-the-pace-takes-exponentially-more-resources\" rel=\"nofollow noopener\" target=\"_blank\">faster chips, bigger data centers, more training data, and greater investment<\/a>. Those can keep driving rapid progress even if the loop can\u2019t sustain itself.<\/p>\n<p>The loop itself could strengthen too. Better training data, memory, and <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7ai-still-struggles-with-big-research-ideas\" rel=\"nofollow noopener\" target=\"_blank\">research judgment<\/a> could all help.<\/p>\n<p>A breakthrough on the scale of the <a href=\"https:\/\/arxiv.org\/abs\/1706.03762\" rel=\"nofollow noopener\" target=\"_blank\">Transformer architecture in 2017<\/a> could change the picture much more. That would be a good reason to revisit these estimates.<\/p>\n<p>Better researchers might also run fewer experiments and learn more from each one. A handful of better ideas can matter more than a mountain of routine runs.<\/p>\n<p>Still, diminishing returns in machine learning aren\u2019t new. <a href=\"https:\/\/papers.nips.cc\/paper_files\/paper\/1993\/file\/1aa48fc4880bb0c9b8a3bf979d3b917e-Paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Cortes and colleagues<\/a> were <a href=\"https:\/\/papers.nips.cc\/paper_files\/paper\/1993\/file\/1aa48fc4880bb0c9b8a3bf979d3b917e-Paper.pdf\" rel=\"nofollow noopener\" target=\"_blank\">fitting machine learning scaling curves in 1993<\/a>: More examples reduced error, following a power law with diminishing returns. These diminishing returns and harsh scaling laws are as old as machine learning. They didn\u2019t appear for the first time with transformers or LLMs or deep learning. That doesn\u2019t prove today\u2019s relationships will last forever. But until we see evidence that we\u2019ve found a new approach that scales without these inhibitors, we should plan for diminishing returns as likely to be with us for some time.<\/p>\n<p>That said, the world is more than just software. <a href=\"https:\/\/basilhalperin.com\/papers\/singularities.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek<\/a> model software progress, hardware progress, and economic feedback together. Better AI helps design better chips; better chips support better AI; economic growth finances more investment in both. Several feedback loops can combine to overcome diminishing returns even when one loop alone can\u2019t. I think it\u2019s fantastic that someone has attempted a model that integrates all these different avenues of improving AI through software, hardware, and economics.<\/p>\n<p>But I have questions about the software loop itself. In their central calibration, fully automating software research puts that loop roughly at the threshold for explosive growth, even without help from better hardware or broader economic growth. Recall that Cunningham\u2019s model puts the self-sustaining threshold at roughly 15% more research productivity per additional ECI point, while our estimate using OpenAI\u2019s experimental data puts today\u2019s gains at only 2-3%. These models use different measures, so we can\u2019t equate their numbers directly. But the contrast matters: their fully automated software loop reaches the threshold, while our best estimate from current data puts today\u2019s loop far below it.<\/p>\n<p>Having AI do all the research doesn\u2019t eliminate the diminishing returns inherent to improving AI, or the broader problem of useful ideas getting harder to find. This is the distinction between Type 4 and Type 5 in the taxonomy above. An AI might autonomously design, train, and test its successor, and still need exponentially more resources to make each additional step forward. Closing the loop doesn\u2019t tell us whether it\u2019s strong enough to sustain itself.<\/p>\n<p>The authors do account for diminishing returns. The concern is whether their calibration overestimates how much useful AI research each round of software improvement produces. Diminishing returns appear to be fundamental to machine learning. We see them in training, in test-time compute, and in the search for better algorithms. Full autonomy could remove human bottlenecks without removing any of those constraints.<\/p>\n<p>We\u2019ve already seen this within autonomous research. In the <a href=\"https:\/\/github.com\/karpathy\/autoresearch\" rel=\"nofollow noopener\" target=\"_blank\">Karpathy autoresearch<\/a> example above, most of the gains arrived early, and more experiments bought progressively less improvement. That was a small experiment with a fixed teacher model, not a test of fully autonomous RSI. It doesn\u2019t settle the question. But it illustrates why removing the human from an experiment loop doesn\u2019t, by itself, remove diminishing returns.<\/p>\n<p>I do expect the feedback loop to get stronger over time. Better AI should become better at research. But based on our best current data, reaching self-sustaining feedback requires a loop roughly five to ten times stronger than today\u2019s. Treating fully automated software research as already at that threshold is a substantial leap, before we add the benefits of hardware improvements or economic growth. I could be wrong, but I\u2019d like to see evidence that autonomy brings enough additional useful discoveries to close that gap.<\/p>\n<p>On hardware, I have some further reservations. The model doesn\u2019t explicitly include the years it can take to turn a chip design into deployed hardware. The authors discuss physical bottlenecks, and I\u2019d like to see manufacturing and construction delays built into the predictions.<\/p>\n<p>I also wonder how much past chip progress came from better ideas, and how much depended on ever more expensive factories and equipment. If we give researchers too much credit for gains that also needed those investments, we could overestimate what faster AI research alone would produce.<\/p>\n<p>Even with those reservations, this is the most compelling paper and model I\u2019ve seen for combining feedback loops in software, hardware, and economics to understand how fast they could push AI forward. I\u2019m not convinced it establishes that a fast AI takeoff is possible under realistic conditions. More data could help us calibrate that judgment. But it gives us a useful framework for understanding what could happen beyond the software layer alone.<\/p>\n<p>This is an important paper that helps us model AI as part of a broader economy that might have larger feedback loops around it. I appreciate it, and I\u2019m glad they wrote it.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p>These estimates rest on less data than I\u2019d like. I might be putting too much weight on a few observations and reaching a comforting conclusion I want to believe. We need better measurements, shared often enough to catch changes as they happen.<\/p>\n<p>When OpenAI released its research data, <a href=\"https:\/\/cherylwu3.github.io\/\" rel=\"nofollow noopener\" target=\"_blank\">Cheryl Wu<\/a> <a href=\"https:\/\/x.com\/cherylwoooo\/status\/2096741494866715055\" rel=\"nofollow\">welcomed the disclosure and pointed out how much was still missing<\/a>. More tokens and experiments are useful things to know about. We also need to see how they turn into better algorithms and more capable AI.<\/p>\n<p><a href=\"https:\/\/x.com\/cherylwoooo\/status\/2096741494866715055\" target=\"_blank\" rel=\"noopener noreferrer nofollow\" data-component-name=\"Twitter2ToDOM\" class=\"pencraft pc-display-contents pc-reset\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/pbs.substack.com\/profile_images\/1728532460131274752\/raUIZ4N0.jpg\" alt=\"X avatar for @cherylwoooo\"  width=\"40\" height=\"40\" draggable=\"false\" loading=\"lazy\" class=\"img-OACg1c object-fit-cover-u4ReeV pencraft pc-reset\"\/><\/p>\n<p>Cheryl Wu@cherylwoooo<\/p>\n<p>I appreciate this as an initial step toward more transparent reporting on RSI. But there is still much further data we need to fully understand RSI.<\/p>\n<p>In particular, OAI disclosed some evidence about the inference compute usage, number of experiments\/researcher, and to a lesser\u2026<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/HRkem5xboAALQAx.jpg\" loading=\"lazy\" class=\"image-c_FmAR\"\/><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/pbs.substack.com\/profile_images\/2007992653586333696\/ViMSkJY0.jpg\" alt=\"X avatar for @kliu128\"  width=\"20\" height=\"20\" draggable=\"false\" class=\"img-OACg1c object-fit-cover-u4ReeV pencraft pc-reset\"\/><\/p>\n<p>Kevin Liu @kliu128<\/p>\n<p>Today we&#8217;re releasing data on models accelerating research at OpenAI. <\/p>\n<p>Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more<\/p>\n<p>11:25 PM \u00b7 Sep 6, 2026 \u00b7 19.7K Views<\/p>\n<p>9 Replies \u00b7 16 Reposts \u00b7 121 Likes<\/p>\n<p><\/a><\/p>\n<p>Figure 32. Cheryl Wu on OpenAI\u2019s research data. <a href=\"https:\/\/x.com\/cherylwoooo\/status\/2096741494866715055\" rel=\"nofollow\">Source<\/a>.<\/p>\n<p>Now Wu, Arjun Ramani, and Basil Halperin, with their colleagues at the <a href=\"https:\/\/elasticity.institute\/\" rel=\"nofollow noopener\" target=\"_blank\">Elasticity Institute<\/a>, have written a concrete proposal: <a href=\"https:\/\/elasticity.institute\/how-to-measure-rsi.pdf\" rel=\"nofollow noopener\" target=\"_blank\">How to Measure RSI<\/a>. It lists eight things the labs could share to help answer these questions. Check it out.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!pB3g!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85097256-48f0-4adb-9183-cdaae1fd600d_3600x1634.jpeg\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/85097256-48f0-4adb-9183-cdaae1fd600d_3600.jpeg\" width=\"1456\" height=\"661\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/85097256-48f0-4adb-9183-cdaae1fd600d_3600x1634.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:661,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:945434,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image\/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.rameznaam.com\/i\/217505960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85097256-48f0-4adb-9183-cdaae1fd600d_3600x1634.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\" title=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Figure 33. Eight proposals for measuring RSI. <a href=\"https:\/\/elasticity.institute\/how-to-measure-rsi.pdf\" rel=\"nofollow noopener\" target=\"_blank\">Source<\/a>.<\/p>\n<p>I\u2019d especially like to see how much useful research each new model adds, holding resources roughly constant, and how that research translates into better AI. That\u2019s how we\u2019ll learn whether the loop is getting stronger.<\/p>\n<p>AI is already helping build better AI. It\u2019s improving at a stupendous pace, and I expect that to continue. We already have <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7we-already-have-narrow-superintelligence\" rel=\"nofollow noopener\" target=\"_blank\">narrow superintelligence<\/a> in chess and Go. I expect increasingly superhuman performance in parts of formal math, coding, and cybersecurity, and any other verifiable domain where machines can generate training data and verify success at machine speed. Those are powerful capabilities. That doesn\u2019t mean we\u2019re close to super-intelligence for less verifiable, messier, open-ended work &#8211; or to a general ASI.<\/p>\n<p>I\u2019m skeptical of a fast takeoff to super-intelligence, but evidence matters more than hunches. Let\u2019s collect the data we need to get a clearer picture of what\u2019s happening. Including evidence that could change our minds. If better AI starts producing <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7how-fast-are-gains-coming-now\" rel=\"nofollow noopener\" target=\"_blank\">enough useful research<\/a> to make the next round easier, I want to know. If the <a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A74-the-sharp-diminishing-returns-to-impressive-ai-numbers\" rel=\"nofollow noopener\" target=\"_blank\">gains keep shrinking<\/a>, I want to know that too.<\/p>\n<p><a href=\"https:\/\/www.rameznaam.com\/p\/471bbae4-1163-4048-944b-18f8b0bf0e41#%C2%A7contents\" rel=\"nofollow noopener\" target=\"_blank\">Back to contents<\/a><\/p>\n<p data-attrs=\"{&quot;url&quot;:&quot;https:\/\/www.noahpinion.blog\/p\/wheres-the-intelligence-explosion?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}\" data-component-name=\"ButtonCreateButton\" class=\"button-wrapper\"><a href=\"https:\/\/www.noahpinion.blog\/p\/wheres-the-intelligence-explosion?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share\" class=\"button primary\" rel=\"nofollow noopener\" target=\"_blank\">Share<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"Art by GPT-6 One of my fundamental beliefs about the world is that Ramez Naam ought to blog&hellip;\n","protected":false},"author":2,"featured_media":785427,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-785426","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/785426","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=785426"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/785426\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/785427"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=785426"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=785426"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=785426"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}