{"id":781864,"date":"2026-09-25T05:38:14","date_gmt":"2026-09-25T05:38:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/781864\/"},"modified":"2026-09-25T05:38:14","modified_gmt":"2026-09-25T05:38:14","slug":"the-specter-of-neuralese-by-scott-alexander","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/781864\/","title":{"rendered":"The Specter Of Neuralese &#8211; by Scott Alexander"},"content":{"rendered":"<p>In the public debates after the release of <a href=\"https:\/\/ai-2027.com\/\" rel=\"nofollow noopener\" target=\"_blank\">AI 2027<\/a>, some people accused us of misleading readers by positing \u201c<a href=\"https:\/\/ai-2027.com\/#narrative-2027-03-31\" rel=\"nofollow noopener\" target=\"_blank\">neuralese recurrence<\/a>\u201d, a hypothetical technology that would allow AIs to think dangerous thoughts without getting detected.<\/p>\n<p>You\u2019ll never guess what happened earlier this month!<\/p>\n<p>But what is recurrence? Is this exactly the same as the neuralese recurrence in AI 2027? And how worrying is it?<\/p>\n<p>The transformer, the technology behind most modern AI, contains some number of layers (modern frontier models are probably around 100). By convention, the input layer is called the \u201cbottom\u201d and the output layer the \u201ctop\u201d. When a transformer does next-word prediction, it takes the last word in at the bottom layer, spends the middle layers processing it, and outputs the predicted next word at the \u201ctop\u201d.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!LCt7!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1c56ad0-f9fd-4d1e-a341-97da427201bb_278x289.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a1c56ad0-f9fd-4d1e-a341-97da427201bb_278x.png\" width=\"278\" height=\"289\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a1c56ad0-f9fd-4d1e-a341-97da427201bb_278x289.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:289,&quot;width&quot;:278,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7320,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa1c56ad0-f9fd-4d1e-a341-97da427201bb_278x289.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   class=\"sizing-normal\"\/><\/a><\/p>\n<p>What if it\u2019s working on a hard question that needs more than 100 layers of processing? From 2017 &#8211; 2024, the answer was \u201cyou\u2019re screwed\u201d. From 2024 onward, the answer was: it outputs its intermediate result, as text, onto a scratchpad called the \u201cchain-of-thought\u201d. Then it runs the transformer again on the intermediate result. Then it repeats until it has a final result.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!cJfR!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4005332-f4fc-41b1-b8fe-9705759a7e3d_892x332.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a4005332-f4fc-41b1-b8fe-9705759a7e3d_892x.png\" width=\"716\" height=\"266.4932735426009\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/a4005332-f4fc-41b1-b8fe-9705759a7e3d_892x332.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:332,&quot;width&quot;:892,&quot;resizeWidth&quot;:716,&quot;bytes&quot;:15003,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4005332-f4fc-41b1-b8fe-9705759a7e3d_892x332.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a>This shows three cycles, but in reality there can be thousands.<\/p>\n<p>(Why does it need a scratchpad, instead of feeding its intermediate results directly into user output and then reading those back into itself? First, the intermediate results may be hundreds of pages long, and the user doesn\u2019t want to read those. Second, the AI companies want to train the AIs to do the intermediate thinking properly, and this benefits from a \u2018psychological\u2019 distinction between intermediate pondering and final user output.)<\/p>\n<p>This was purely a capabilities play &#8211; models are smarter when they can keep thinking instead of limiting themselves to 100 processing steps &#8211; but it coincidentally was very good for safety. The chain-of-thought scratchpad is written in English (although some Chinese models use an English-Chinese hybrid, and other AIs develop their own weird jargon). You can just read what the AI is thinking! If the AI is thinking \u201cBetter hack some websites, then kill all humans\u201d, you can shut it down. Maybe not actually &#8211; if you have thousands of AIs writing millions of pages of scratchpad, you can\u2019t read all of that in real-time, and will need to delegate the task to fallible AI monitors. But in theory this ought to work.<\/p>\n<p>But language is slower and shallower than thought. For 100 steps in a row, the AI can shoot delicate subtle ideas from layer to layer at the speed of light. Then there\u2019s one step where it has to encode them into twenty-six glyphs invented by Phoenician turquoise miners in 1800 BC. Then it has to re-encode the Phoenician glyphs into delicate subtle lightspeed ideas before it can do anything else. This has long been acknowledged as a bottleneck in existing transformers. So: what if they could finish their 100 layers of thinking, then send the resulting thought back to the first layer for more processing, rather than sending a text scratchpad on which they had written the thought?<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!wcju!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd90963f0-1bc1-4922-8927-a2b50bedfc6d_332x332.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/d90963f0-1bc1-4922-8927-a2b50bedfc6d_332x.png\" width=\"332\" height=\"332\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/d90963f0-1bc1-4922-8927-a2b50bedfc6d_332x332.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:332,&quot;width&quot;:332,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:7085,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd90963f0-1bc1-4922-8927-a2b50bedfc6d_332x332.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>This is neuralese recurrence &#8211; \u201crecurrence\u201d because it\u2019s going back in a loop, \u201cneuralese\u201d because the thing that\u2019s looping is the thought itself, in the native language of thought, rather than words.<\/p>\n<p>The AI\u2019s \u201clanguage of thought\u201d looks like a vector, thousands of numbers long. An \u201cintermediate result\u201d in this scheme might look something like (0.4, 0.1, 5, 0.443, \u2026 and so on for thousands of numbers). We don\u2019t know how to read these. The science of reading these thoughts is a subfield of <a href=\"https:\/\/www.astralcodexten.com\/p\/god-help-us-lets-try-to-learn-about\" rel=\"nofollow noopener\" target=\"_blank\">AI interpretability, which is still in its infancy<\/a>. If an AI with neuralese recurrence were to think \u201cBetter hack some websites, then kill all humans\u201d, it would look like (0.4, 0.1, 5, 0.443, \u2026 and so on for thousands of numbers), and we would never find out. This is why AI 2027 discussed it as a prelude to AIs that could escape human monitoring.<\/p>\n<p>So did OpenAI achieve this? Their chief scientist, Jakub Pachocki, says \u201cnot really\u201d:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/x.com\/merettm\/status\/2095023204993490967\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/44c364f9-f3b9-4129-a607-d60b2536ad74_582x.png\" width=\"582\" height=\"322\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/44c364f9-f3b9-4129-a607-d60b2536ad74_582x322.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:322,&quot;width&quot;:582,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:34907,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:&quot;https:\/\/x.com\/merettm\/status\/2095023204993490967&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44c364f9-f3b9-4129-a607-d60b2536ad74_582x322.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>What does Pachocki mean by \u201cthe depth of the computation graph . . . is within a factor of two of GPT-4\u201d? And should this resolve alignment concerns?<\/p>\n<p>Our example transformer has four green layers: a layer depth of four:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!p-qB!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c3ba660-916a-4a11-98b8-401ffb542917_383x392.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/9c3ba660-916a-4a11-98b8-401ffb542917_383x.jpeg\" width=\"341\" height=\"349.0130548302872\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/9c3ba660-916a-4a11-98b8-401ffb542917_383x392.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:392,&quot;width&quot;:383,&quot;resizeWidth&quot;:341,&quot;bytes&quot;:30042,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c3ba660-916a-4a11-98b8-401ffb542917_383x392.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Now we add a little recurrence: instead of layer 3 going to layer 4, it loops back to layer 2, and the data go through this loop three times, then finally progress to layer 4:<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!dP9n!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcdfac3c-bfb3-4c21-9ae8-32caf2516f12_320x332.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/dcdfac3c-bfb3-4c21-9ae8-32caf2516f12_320x.png\" width=\"320\" height=\"332\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/dcdfac3c-bfb3-4c21-9ae8-32caf2516f12_320x332.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:332,&quot;width&quot;:320,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6709,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdcdfac3c-bfb3-4c21-9ae8-32caf2516f12_320x332.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>Here the data are getting eight layers of processing: 1 \u2192 2 \u2192 3 \u2192 2 \u2192 3 \u2192 2 \u2192 3 \u2192 4. So this recurrent transformer has a depth of eight.<\/p>\n<p>In this sense, \u201crecurrence\u201d, despite sounding scary, is no worse than just building an eight-layer transformer to begin with, which was always an option.<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!j_gj!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a82499d-8b81-463e-955f-9b2816251eb7_508x587.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/8a82499d-8b81-463e-955f-9b2816251eb7_508x.png\" width=\"428\" height=\"494.5590551181102\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/8a82499d-8b81-463e-955f-9b2816251eb7_508x587.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:587,&quot;width&quot;:508,&quot;resizeWidth&quot;:428,&quot;bytes&quot;:13906,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a82499d-8b81-463e-955f-9b2816251eb7_508x587.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>That is, from an alignment perspective, suppose it takes 1,000 layers\u2019 worth of processing to come up with an evil plot. A normal AI can think for 100 layers, but then it has to record its output on the chain-of-thought &#8211; something like \u201cHere\u2019s one-tenth of the evil plot I\u2019m trying to hatch\u201d &#8211; and we can catch it. So we want to avoid letting the model compute for 1,000 layers straight without using the scratchpad. This means don\u2019t create a recurrent model that can loop 100 layers 10x, but it also means don\u2019t build a 1,000 layer model to begin with.<\/p>\n<p>So when Pachocki says that Astra only has twice the depth of GPT-4, he\u2019s saying something like: look, guys, we\u2019ve been adding layers to our AIs for years. GPT-2 had 48 layers, GPT-3 had 96, and we didn\u2019t tell you how many GPT-4 had but let\u2019s say it was 120. At none of these points did you complain, because it wasn\u2019t \u201crecurrence\u201d, just a normal natural 120 layers. Now Astra is &#8211; let\u2019s say &#8211; 180 layers. It\u2019s true that we got these extra layers by recurrence rather than literally building a bigger transformer, but the alignment implications are no different. For all you know, Anthropic built a literal 180 layer transformer yesterday, and you guys didn\u2019t bother them. The important thing is that we not increase our layer number by some crazy amount, and we didn\u2019t. We\u2019re just adding layers the same way lots of other AI architectures did, albeit by other means.<\/p>\n<p>How convincing is this argument?<\/p>\n<p>When reading this, my question was &#8211; why don\u2019t you just loop all the layers one million times, and get an AI with a million times the computational depth of the shmuck who didn\u2019t do that? Isn\u2019t this free capabilities?<\/p>\n<p>A looped transformer model like Astra uses loops as a hack to add more layers. This hack isn\u2019t as good as adding more layers for real. A real new layer gives the AI more \u201croom\u201d to store knowledge and thought styles. A looped layer just lets the AI use the same knowledge and thought style more times.<\/p>\n<p>This is far from useless. A mathematician might spend years contemplating the same problem before getting it right; all through this period, he is using the same type of thinking (the mathematical knowledge and techniques his brain is capable of). But sometimes doing it for a year is better than doing it for a minute.<\/p>\n<p>But it\u2019s not optimal either. An AI with more real layers can actually be \u201csmarter\u201d in the conventional sense of the term. Even though it might help to give the same mathematician more time to think about the problem, holding amount of time constant, it will help even more to give the problem to a smarter mathematician with more training.<\/p>\n<p>I asked Fable to estimate the relative capabilities gain, measured in \u201cgenerations\u201d (eg GPT-4 to GPT-5) from three things:<\/p>\n<p>Using chain-of-thought (baseline)<\/p>\n<p>Doubling the \u201csimulated\u201d number of layers by a recurrent loop.<\/p>\n<p>Doubling the \u201creal\u201d number of layers.<\/p>\n<p>Its answers were 0, +0.03, and +0.25, with wide variation depending on the type of task. Its headline result was that the gains from extra simulated layers are much smaller than the gains from extra real layers.<\/p>\n<p>But why are we merely doubling the number of layers? Once you can loop layers at all, why not loop them a million times? Pachocki says he\u2019s replacing something like 120 layers \u2192 chain-of-thought scratchpad \u2192 another 120 layers with something more like 240 layers \u2192 chain-of-thought scratchpad \u2192 another 240 layers, but why doesn\u2019t he just loop the AI endlessly until it comes up with a final result?<\/p>\n<p>Here is a diagram of a transformer process (<a href=\"https:\/\/arxiv.org\/abs\/2507.11473\" rel=\"nofollow noopener\" target=\"_blank\">source<\/a>):<\/p>\n<p><a target=\"_blank\" href=\"https:\/\/substackcdn.com\/image\/fetch\/$s_!wLdk!,f_auto,q_auto:good,fl_progressive:steep\/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff923b3a1-e3c3-4091-a31d-38e841c34a7d_625x378.png\" data-component-name=\"Image2ToDOM\" class=\"image-link image2 is-viewable-img can-restack\" rel=\"nofollow noopener\"><img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2026\/09\/https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f923b3a1-e3c3-4091-a31d-38e841c34a7d_625x.jpeg\" width=\"625\" height=\"378\" data-attrs=\"{&quot;src&quot;:&quot;https:\/\/substack-post-media.s3.amazonaws.com\/public\/images\/f923b3a1-e3c3-4091-a31d-38e841c34a7d_625x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:625,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:64977,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image\/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https:\/\/www.astralcodexten.com\/i\/215243350?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff923b3a1-e3c3-4091-a31d-38e841c34a7d_625x378.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" alt=\"\"   loading=\"lazy\" class=\"sizing-normal\"\/><\/a><\/p>\n<p>On each token (for example, the word \u201cso\u201d) the transformer looks at its existing context and runs a forward pass (the gray up arrows) through some number of layers. Then it finishes its forward pass, emits a token of chain-of-thought (the curvy blue arrows), then starts another forward pass on the new context (the old context plus the new token).<\/p>\n<p>Adding more layers lets the AI think more between chain-of-thought tokens, improving the quality of each new token. But it uses extra compute. If the user is on a budget, the more compute that the AI spends on each chain-of-thought token, the fewer chain-of-thought tokens it can emit. Holding compute fixed, there\u2019s some optimal balance between the quality and quantity of chain-of-thought tokens, and OpenAI must have decided it lay at looping 2-4x but not 1000x.<\/p>\n<p>But why use chain-of-thought at all? Why not just loop unboundedly, until the problem is complete?<\/p>\n<p>(on the diagram above, that would look like replacing the English-reasoning blue arrows with neuralese-reasoning gray arrows)<\/p>\n<p>This would work if the AI could be trained to do it, but currently it can\u2019t. Transformers are pre-trained on human text. They know how to use chain-of-thought partly from reading real humans\u2019 chains-of-thought: \u201cHmmmm, this is a hard problem, but it looks sort of like something I\u2019ve solved with differential equations before. Maybe if I plug in a differential equation there . . . no, that would be too inelegant . . wait, what if -\u201d, and partly from reading enough other human output that this kind of human-style thinking comes naturally. But there is no convenient Internet text about human reasoning entirely in vectors, so the AIs can\u2019t be trained to do it properly. In principle this isn\u2019t fatal \u2013 all AI reasoning is in vectors, and the whole point of training is to teach them to do it well \u2013 but training an entire vector-based thought process for thousands of steps without ever bottoming out in a human-imitating chain-of-thought-token intermediate is prohibitively costly. The researchers who design transformers have to make guesses about what chain-of-thought frequency will maximize capabilities while minimizing training cost. This is part of what Pachocki was denying. AI 2027\u2019s \u201cneuralese\u201d is the thing researchers currently don\u2019t know how to do \u2013 train an AI to loop as often as it wants, avoiding English chain-of-thought entirely. Pachocki was saying: don\u2019t worry, we still don\u2019t know how to do that.<\/p>\n<p>But now we can see that this is not entirely reassuring. If it takes an AI 1,000 unmonitored gray-arrow steps to devise a dangerous plot, it can get those steps either by replacing all the blue arrows with gray arrows (in which case all its steps are unmonitored, and it can plot as much as it wants), or by having so many layers that there are 1,000 steps in a single vertical forward pass in between blue arrows. It would be prohibitive to make that many real layers, but it\u2019s possible by looping. So (the safety community argues) even Pachocki\u2019s looped transformer architecture is a step along that path.<\/p>\n<p>All of this is starting to get confusing, but the takeaway is:<\/p>\n<p>An AI with a limited number of layers can\u2019t plot at all without having to record the plot on its chain-of-thought.<\/p>\n<p>An AI that uses loops to simulate a large number of layers (\u201clooped transformer\u201d) can make a plot within a single forward pass, in between chain-of-thought steps.<\/p>\n<p>An AI that no longer uses chain-of-thought at all (\u201ctrue neuralese\u201d) can plot at leisure, considering it for as many episodes as it needs, and we will never know.<\/p>\n<p>Every previous transformer has been (1). Astra is getting part of the way to (2). The dangerous AIs in AI 2027 are (3), but so far nobody has invented this in real life.<\/p>\n<p>Having finished that digression, let\u2019s return to the original question. Pachocki says that Astra\u2019s recurrence barely gives it any more layers than the competition. Does that exonerate him from the charge of creating dangerous recurrent AIs?<\/p>\n<p>Here the best thing I\u2019ve read is Linchuan Zhang\u2019s <a href=\"https:\/\/www.lesswrong.com\/posts\/xPkmfsZ3qx4nrAco7\/categorical-taboos-are-much-better-than-threshold-taboos\" rel=\"nofollow noopener\" target=\"_blank\">Categorical Taboos Are Much Better Than Threshold Taboos: Neuralese Edition<\/a>.<\/p>\n<p>Linch says: there\u2019s an emerging taboo on creating the dangerous sort of neuralese AIs seen in AI 2027. Everyone in this discussion agrees that the taboo is correct. The only question is whether OpenAI violated that taboo.<\/p>\n<p>But taboos only work when there are clear boundaries defining what is versus isn\u2019t within the tabooed area. For example, it\u2019s illegal in America to buy alcohol before age 21. Suppose that someone buys alcohol two days before their twenty-first birthday. Should the police arrest them? Obviously this age gap doesn\u2019t really make a difference. We\u2019re using age as a proxy for something like maturity, and there\u2019s so much variation in maturity that there\u2019s no real difference between a 20.99 year old and a 21 year old. Still, although the police might sometimes choose to overlook this, we have to at least maintain the fiction that this is an arrestable offense. Why? Suppose that we made a specific boundary &#8211; if you\u2019re above 20.8 years old, we won\u2019t arrest you. But then we\u2019ve just decreased our bright line from 21 to 20.8. And we could encounter the same problem: a 20.799-year-old buys alcohol and the police have to decide whether to make an exception or not. If we\u2019re going to draw the line somewhere, it might as well be at the place we\u2019d already promised to draw it, told everybody that we\u2019d drawn it, etc.<\/p>\n<p>(Linch\u2019s own example is nuclear weapons. There\u2019s a taboo on using nukes in war. If someone uses a tiny nuke, which produces an explosion no bigger than a conventional bomb, then in some sense this doesn\u2019t matter, because it\u2019s no worse than a conventional bomb would have been. But in another sense, it matters a lot, because you\u2019ve broken the taboo, and now there\u2019s only a weaker, fuzzier taboo preventing you from using a very big nuke.)<\/p>\n<p>So if we want a taboo on recurrence, there needs to be some specific taboo.  When Pachocki says \u201cYes, we did recurrence, but it only added a couple of layers, so it doesn\u2019t really matter,\u201d this is analogous to the case where a drinker tells the police \u201cYes, I\u2019m below 21, but only by a few months, so it doesn\u2019t really matter\u201d, or a despot tells the UN \u201cYes, I used a nuke, but it was very small.\u201d It\u2019s true that it doesn\u2019t matter in real life, but it matters a lot for whether you can maintain a taboo or not.<\/p>\n<p>Unfortunately, we currently lack agreement on what, if any, taboo exists. Some people argue that there should have been a taboo on looping layers (which OpenAI would have broken), and other people say that they\u2019re making that up and that was never a taboo. There\u2019s probably an informal taboo on true neuralese with no chain of thought at all, but it\u2019s a complicated technology with multiple moving parts, there\u2019s no agreement as to which moving part the taboo is on, and without that agreement, some of the parts might slip through the cracks. So what are our options?<\/p>\n<p>First, contra Pachocki, it might be worth instituting a taboo on looping (adding \u201csimulated layers\u201d) at all, even without any taboo on adding more \u201creal layers\u201d. Adding real layers is expensive and generally not worth it; AI companies have only chosen to increase real depth of their models by ~20% per year; at that rate, it might take decades to reach a danger point. Adding \u201csimulated layers\u201d, which are much cheaper and can be jammed in hundreds at a time, is more dangerous. So maybe we should ignore the hypothetical real layer \/ simulated layer equivalence in favor of saying that looping is always taboo, regardless of how many real layers other people are adding somewhere else.<\/p>\n<p>Second, companies could agree on some maximum number of layers (<a href=\"https:\/\/institute.deepmind.com\/essays\/the-case-for-reasoning-transparency\/\" rel=\"nofollow noopener\" target=\"_blank\">real or simulated<\/a>), like 1,000. This probably wouldn\u2019t affect real layers (by the argument above), but it would limit loops to some maximum size.<\/p>\n<p>Third, researchers could hash out what the components of \u201ctrue neuralese\u201d are, and agree not to do them. This would require some conceptual foundations, but it would be the most durable and the closest to closing off the specific dangerous technology that AI 2027 was worried about. <\/p>\n<p>Unlike some other vague attempts to \u201cban superintelligent AI\u201d or \u201cban recursive self-improvement\u201d, these taboos might stick even without strong government action: at least for now, the capability gains from breaking them seem modest, and the dangers are particularly obvious, so they might be maintainable by voluntary commitments even while a \u201crace dynamic\u201d was still going on.<\/p>\n<p>AI company leaders have recently put aside their differences and agreed to \u201c<a href=\"https:\/\/darioamodei.com\/post\/we-must-pace-the-frontier\" rel=\"nofollow noopener\" target=\"_blank\">pace<\/a> <a href=\"https:\/\/x.com\/sama\/status\/2098811563415150910\" rel=\"nofollow\">the <\/a><a href=\"https:\/\/x.com\/elonmusk\/status\/2098789109980332057\" rel=\"nofollow\">frontier<\/a>\u201d in various senses to be fully defined later. Their meetings will probably have a very long list of discussion items, but one more useful thing they could do would be to formalize this taboo, so we know exactly what it is we\u2019re trying to stay away from.<\/p>\n<p>If you\u2019re interested in this topic, please also read <a href=\"https:\/\/blog.redwoodresearch.org\/p\/latent-reasoning-architectures-would\" rel=\"nofollow noopener\" target=\"_blank\">this essay by Redwood Research<\/a>, which explains some of the relevant concepts and technologies in more detail.<\/p>\n","protected":false},"excerpt":{"rendered":"In the public debates after the release of AI 2027, some people accused us of misleading readers by&hellip;\n","protected":false},"author":2,"featured_media":781865,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,86,56,54,55],"class_list":["post-781864","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/781864","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=781864"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/781864\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/781865"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=781864"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=781864"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=781864"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}