{"id":481482,"date":"2026-03-18T03:23:17","date_gmt":"2026-03-18T03:23:17","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/481482\/"},"modified":"2026-03-18T03:23:17","modified_gmt":"2026-03-18T03:23:17","slug":"the-karpathy-loop-700-experiments-2-days-and-a-glimpse-of-where-ai-is-heading","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/481482\/","title":{"rendered":"&#8216;The Karpathy Loop&#8217;: 700 experiments, 2 days, and a glimpse of where AI is heading"},"content":{"rendered":"<p>Earlier this month, Andrej Karpathy, a well-known AI researcher who was one of the founding employees of OpenAI and later headed up AI for <a aria-label=\"Go to https:\/\/fortune.com\/company\/tesla\/\" target=\"_blank\" href=\"https:\/\/fortune.com\/company\/tesla\/\" rel=\"nofollow noopener\">Tesla<\/a>, went viral on <a aria-label=\"Go to https:\/\/fortune.com\/company\/twitter\/\" target=\"_blank\" href=\"https:\/\/fortune.com\/company\/twitter\/\" rel=\"nofollow noopener\">X<\/a>. This alone isn\u2019t so unusual. Karpathy\u2014who now works as an independent AI researcher and is also the founder of Eureka Labs, which says it is creating a new kind of school for the AI era\u2014has 1.9 million followers on X and his reputation is such that almost anything he says about AI is treated as either gospel or prophecy.<\/p>\n<p>But <a aria-label=\"Go to https:\/\/x.com\/karpathy\/status\/2030371219518931079?s=20\" href=\"https:\/\/x.com\/karpathy\/status\/2030371219518931079?s=20\" rel=\"nofollow\">this post<\/a> was about an experiment he\u2019d run where put an AI coding agent to work running a series of experiments to figure out how to improve the training of a small language model. He let the AI agent run continuously for two days, during which time it conducted 700 different experiments. Over the course of those experiments, it discovered 20 optimizations that improved the training time. <\/p>\n<p>Karpathy found that applying the same 20 tweaks to a larger, but still fairly small, language model resulted in an 11% speed up in the time it took to train the model. Karpathy called the system he built for conducting this experiment \u201cautoresearch.\u201d<\/p>\n<p>Tobias L\u00fctke, the cofounder and CEO of <a aria-label=\"Go to https:\/\/fortune.com\/company\/shopify\/\" target=\"_blank\" href=\"https:\/\/fortune.com\/company\/shopify\/\" rel=\"nofollow noopener\">Shopify<\/a>, posted on X that he tried autoresearch to optimize an AI model on internal company data, giving the agent instructions to improve the model\u2019s quality and speed. L\u00fctke reported that after letting autoresearch run overnight, it ran 37 experiments and delivered a 19% performance gain.<\/p>\n<p>What caught many people\u2019s attention was that the autoresearch is close to the idea of self-improving AI systems that were originally broached in science fiction and that some AI researchers fervently desire and others deeply fear. The concern is that \u201crecursive self-improvement,\u201d where an AI continually optimizes its own code and training in a kind of loop, could lead to what AI safety researchers sometimes call a \u201chard takeoff\u201d or an \u201cintelligence explosion.\u201d In these scenarios, an AI system rapidly improves its own performance, leading it to surpass human cognitive abilities and escape human control.<\/p>\n<p>Karpathy\u2019s experiment wasn\u2019t quite this. The AI agent at the heart of autoresearch set up isn\u2019t refining its own training set up, it\u2019s adjusting the training code and initial neural network settings for a different, much smaller and less sophisticated, AI model. But Karpathy rightly noted that his experiment had big implications for how AI labs will do research going forward, and this might accelerate their progress.<\/p>\n<p>\u201cAll LLM frontier labs will do this. It\u2019s the final boss battle,\u201d Karpathy <a aria-label=\"Go to https:\/\/x.com\/karpathy\/status\/2031135152349524125?s=20\" href=\"https:\/\/x.com\/karpathy\/status\/2031135152349524125?s=20\" rel=\"nofollow\">wrote<\/a> on X. He acknowledged that \u201cit\u2019s a lot more complex at scale of course,\u201d since his autoresearcher only had to worry about adjusting a model and training process that was contained in just 630 lines of Python code, whereas the training codebase of frontier AI models is orders of magnitude bigger. \u201cBut doing it is \u2018just engineering\u2019 and it\u2019s going to work,\u201d he continued. \u201cYou spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.\u201d<\/p>\n<p>He said that while the current autoresearch system he built was designed for a single agent to continually improve a piece of code along a single path, in the future he imagines multiple AI agents will be able to explore different optimizations and different experiments in parallel. \u201cThe next step for autoresearch is that it has to be asynchronously massively collaborative for agents,\u201d he <a aria-label=\"Go to https:\/\/x.com\/karpathy\/status\/2030705271627284816?s=20\" href=\"https:\/\/x.com\/karpathy\/status\/2030705271627284816?s=20\" rel=\"nofollow\">wrote<\/a>. \u201cThe goal is not to emulate a single PhD student, it\u2019s to emulate a research community of them.\u201d<\/p>\n<p>Karpathy also said something else about autoresearch which got many people excited. \u201c*any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm,\u201d he wrote. \u201cIt\u2019s worth thinking about whether your problem falls into this bucket too.\u201d<\/p>\n<p>Some commentators pointed out that the basic components of autoresearch could be used for many other agentic systems to optimize a process. Janakiram MSV, principal analyst at Janakiram &amp; Associates, <a aria-label=\"Go to https:\/\/thenewstack.io\/karpathy-autonomous-experiment-loop\/\" href=\"https:\/\/thenewstack.io\/karpathy-autonomous-experiment-loop\/\" rel=\"nofollow noopener\" target=\"_blank\">writing<\/a> in tech publication The New Stack called this \u201cthe Karpathy Loop.\u201d It has three components: an agent with access to a single file that it can modify; a single metric, objectively testable metric, that the agent can optimize for; and a fixed time limit for how long each experiment can run. He also highlighted that the instructions Karpathy gave the AI agent in autoresearch were also good models for anyone interacting with any AI agent. The plain text file Karpathy used included clear instructions for what the agent should do, constraints, telling the agent what it should not do or change, and a stopping criteria, indicating how long each loop should run and when the agent should stop looping and report its results.<\/p>\n<p>But some critics <a aria-label=\"Go to https:\/\/x.com\/ahatamiz1\/status\/2031228183421530471?s=20\" href=\"https:\/\/x.com\/ahatamiz1\/status\/2031228183421530471?s=20\" rel=\"nofollow\">said<\/a> that Karpathy had done little more than rediscover part of a process known as <a aria-label=\"Go to https:\/\/arxiv.org\/abs\/1908.00709\" href=\"https:\/\/arxiv.org\/abs\/1908.00709\" rel=\"nofollow noopener\" target=\"_blank\">AutoML<\/a> that researchers at <a aria-label=\"Go to https:\/\/fortune.com\/company\/alphabet\/\" target=\"_blank\" href=\"https:\/\/fortune.com\/company\/alphabet\/\" rel=\"nofollow noopener\">Google<\/a>, <a aria-label=\"Go to https:\/\/fortune.com\/company\/microsoft\/\" target=\"_blank\" href=\"https:\/\/fortune.com\/company\/microsoft\/\" rel=\"nofollow noopener\">Microsoft<\/a>, and other AI labs have already been using for years. AutoML also uses an optimization loop and series of experiments to find the best data to use for AI, the best model architecture to use, and to tune that model architecture. But it doesn\u2019t use an AI agent that can read AI research papers and develop hypotheses for which improvement to make. AutoML systems tend to depend on random variations or various evolutionary algorithms to decide which changes to try.\u00a0<\/p>\n<p>Karpathy replied to some of these comments, saying that some AutoML methods, such as neural architecture search, which is an automated way to optimize the design of an AI model, were not nearly as powerful as his autoresearch. \u201cNeural architecture search as it existed then is such a weak version of this that it\u2019s in its own category of totally useless by comparison,\u201d he <a aria-label=\"Go to https:\/\/x.com\/karpathy\/status\/2031138678647783869?s=20\" href=\"https:\/\/x.com\/karpathy\/status\/2031138678647783869?s=20\" rel=\"nofollow\">wrote<\/a>. \u201cThis is an *actual* LLM writing arbitrary code, learning from previous experiments, with access to the internet. It\u2019s not even close.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Earlier this month, Andrej Karpathy, a well-known AI researcher who was one of the founding employees of OpenAI&hellip;\n","protected":false},"author":2,"featured_media":481483,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,10606,733,4308,6112,11376,874,547,33413,90,86,1216,56,54,55],"class_list":["post-481482","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-agents","tag-artificial-intelligence","tag-artificialintelligence","tag-computer-science","tag-machine-learning","tag-openai","tag-research","tag-research-and-development","tag-science","tag-technology","tag-tesla","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/481482","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=481482"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/481482\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/481483"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=481482"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=481482"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=481482"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}