{"id":273737,"date":"2025-11-09T18:16:29","date_gmt":"2025-11-09T18:16:29","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/273737\/"},"modified":"2025-11-09T18:16:29","modified_gmt":"2025-11-09T18:16:29","slug":"ai-turns-brain-scans-into-full-sentences-and-its-eerie-to-say-the-least","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/273737\/","title":{"rendered":"AI Turns Brain Scans Into Full Sentences and It\u2019s Eerie To Say The Least"},"content":{"rendered":"<p><a href=\"https:\/\/cdn.zmescience.com\/wp-content\/uploads\/2025\/11\/Neuroradiology-min-scaled-1.webp\" rel=\"nofollow noopener\" target=\"_blank\"><img src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2025\/11\/Neuroradiology-min-scaled-1-1024x643.webp.webp\" height=\"643\" width=\"1024\"   class=\"wp-image-293522 sp-no-webp\" alt=\"\" fetchpriority=\"high\" decoding=\"async\"\/> <\/a>Credit: Oryon.<\/p>\n<p>In a dark MRI scanner outside Tokyo, a volunteer watches a video of someone hurling themselves off a waterfall. Nearby, a computer digests the brain activity pulsing across millions of neurons. A few moments later, the machine produces a sentence: \u201cA person jumps over a deep water fall on a mountain ridge.\u201d<\/p>\n<p>No one typed those words. No one spoke them. They came directly from the volunteer\u2019s brain activity.<\/p>\n<p>That\u2019s the startling premise of \u201cmind captioning,\u201d a new method developed by Tomoyasu Horikawa and colleagues at NTT Communication Science Laboratories in Japan. Published this week in <a href=\"https:\/\/www.science.org\/doi\/10.1126\/sciadv.adw1464\" rel=\"nofollow noopener\" target=\"_blank\">Science Advances<\/a>, the system uses a blend of brain imaging and artificial intelligence to generate textual descriptions of what people are seeing \u2014 or even visualizing with their mind\u2019s eye \u2014 based only on their neural patterns.<\/p>\n<p>As <a href=\"https:\/\/www.nature.com\/articles\/d41586-025-03624-1\" rel=\"nofollow noopener\" target=\"_blank\">Nature<\/a> journalist Max Kozlov put it, the technique \u201cgenerates descriptive sentences of what a person is seeing or picturing in their mind using a read-out of their brain activity, with impressive accuracy.\u201d<\/p>\n<p>This is not the stuff of science fiction anymore. It\u2019s not mind-reading either, at least not yet. But it\u2019s a vivid demonstration of how our brains and modern AI models might be speaking a surprisingly similar language.<\/p>\n<p>Decoding Meaning from the Silent Mind<\/p>\n<p><a href=\"https:\/\/cdn.zmescience.com\/wp-content\/uploads\/2025\/11\/sciadv.adw1464-f1-scaled.jpg\" rel=\"nofollow noopener\" target=\"_blank\"><img loading=\"lazy\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2025\/11\/sciadv.adw1464-f1-1024x395.jpg\" height=\"395\" width=\"1024\"   class=\"wp-image-293508 sp-no-webp\" alt=\"\" decoding=\"async\"\/> <\/a>The researchers trained an AI to link brain scans with video captions, then used it to turn new brain activity \u2014 whether from watching or recalling scenes \u2014 into sentences through an iterative word-replacement process guided by language models. Credit: Nature, 2025, Horikawa.<\/p>\n<p>To build the system, Horikawa had to bridge two universes: the intricate geometry of human thought and the sprawling semantic web that language models use to understand words. Six volunteers spent nearly seventeen hours each in an MRI scanner, watching 2,180 short, silent video clips. The scenes ranged from playful animals to emotional interactions, abstract animations, and everyday moments. Each clip lasted only a few seconds, but together they provided a massive dataset of how the brain reacts to visual experiences.<\/p>\n<p>For every video, the researchers also gathered twenty captions written by online volunteers. The captions were complete sentences describing what was happening in each scene. The captions were cleaned up with the help of ChatGPT. Each sentence was then transformed into a complex numerical signature \u2014 a point in a vast multi-vector semantic space \u2014 using a language model called DeBERTa.<\/p>\n<p>The team then mapped the brain activity recorded during each video to these semantic signatures. In other words, they trained an AI to recognize what kinds of neural patterns corresponded to particular kinds of meaning. Instead of using deep, opaque neural networks, the researchers relied on a more transparent linear model. This model could reveal which regions of the brain contributed to which kinds of semantic information.<\/p>\n<p>From Abstract Meaning to Words<\/p>\n<p>Once the system could predict the \u201cmeaning vector\u201d of what someone was watching, it faced the next challenge: turning that abstract representation into an actual sentence. To do that, the Japanese scientist used another language model, RoBERTa, to generate words step by step. It began with a meaningless placeholder and, over a hundred iterations, filled in blanks, tested alternative sentences, and kept whichever version best matched the decoded meaning.<\/p>\n<p>The process resembled an evolution of language inside the machine\u2019s circuits. Early attempts sounded like nonsense but with each refinement, the sentences grew more accurate, finally converging on a full, coherent description of the scene.<\/p>\n<p>When tested, the system could match the correct video to its generated description about half the time, even when presented with a hundred possibilities. \u201cThis is hard to do,\u201d Alex Huth, a neuroscientist at the University of California, Berkeley, who has worked on similar brain-decoding projects, told Nature. \u201cIt\u2019s surprising you can get that much detail.\u201d<\/p>\n<p>The researchers also made a surprising discovery when they scrambled the word order of the generated captions. The quality and accuracy dropped sharply, showing that the AI wasn\u2019t just picking up on keywords but grasping something deeper \u2014 perhaps the structure of meaning itself, the relationships between objects, actions, and context.<\/p>\n<p lang=\"en\" dir=\"ltr\">Our new paper is on bioRxiv.<br \/>We present a novel generative decoding method, called Mind Captioning, and demonstrate the generation of descriptive text of viewed and imagined content from human brain activity.<\/p>\n<p>The video shows text generated for viewed content during optimization. <a href=\"https:\/\/t.co\/e0cP6B3CDL\" rel=\"nofollow\">https:\/\/t.co\/e0cP6B3CDL<\/a> <a href=\"https:\/\/t.co\/mB2CO959tT\" rel=\"nofollow\">pic.twitter.com\/mB2CO959tT<\/a><\/p>\n<p>\u2014 Tomoyasu Horikawa (@HKT52) <a href=\"https:\/\/twitter.com\/HKT52\/status\/1784094557229179260?ref_src=twsrc%5Etfw\" rel=\"nofollow noopener\" target=\"_blank\">April 27, 2024<\/a><\/p>\n<p>The Language of Thought<\/p>\n<p>One of the most striking experiments came later, when the volunteers were asked to recall the videos rather than watch them. They closed their eyes, imagined the scenes, and rated how vivid their mental replay felt. The same model, trained only on perception data, was used to decode these recollections. Astonishingly, it still worked.<\/p>\n<p><a href=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2025\/11\/image-11.png\"><img loading=\"lazy\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2025\/11\/image-11.png\" height=\"754\" width=\"979\"   class=\"wp-image-293510 sp-no-webp\" alt=\"Image illustrating images and text from the study\" decoding=\"async\"\/> <\/a> Credit: Nature, 2025, Horikawa. <\/p>\n<p>Even when subjects were only imagining the videos, the AI generated accurate sentences describing them, sometimes identifying the right clip out of a hundred. That result hinted at a powerful idea: the brain uses similar representations for seeing and visual recall, and those representations can be translated into language without ever engaging the traditional \u201clanguage areas\u201d of the brain.<\/p>\n<p>In fact, when the researchers deliberately excluded regions typically associated with language processing, the system continued to generate coherent text. This suggests that structured meaning \u2014 what scientists call \u201csemantic representation\u201d \u2014 is distributed widely across the brain, not confined to speech-related zones.<\/p>\n<p>That discovery carries enormous implications for people who can\u2019t speak. Individuals with aphasia or neurodegenerative diseases that affect language could, in principle, use such systems to communicate through their nonverbal brain activity. The paper calls this an \u201cinterpretive interface\u201d that could restore communication for those whose words are trapped inside their minds.<\/p>\n<p>Promise and Concerns<\/p>\n<p>Still, the researchers are careful not to overpromise. The technology is far from being a mind-reading device. It depends on hours of personalized data from each participant, massive MRI scanners, and a very narrow set of visual stimuli. The sentences it generates are filtered through the biases of the English-language captions and the models used to train them. Change the language model or the dataset, and the output could shift dramatically.<\/p>\n<p>Horikawa himself insists that the system doesn\u2019t reconstruct thoughts directly. It instead translates them through layers of AI interpretation. \u201cTo accurately characterize our primary contribution, it is essential to frame our method as an interpretive interface rather than a literal reconstruction of mental content,\u201d the paper states.<\/p>\n<p>The ethical implications of this technology are hard to ignore. If machines can turn brain activity into words, even imperfectly, who controls that information? Could it be misused in surveillance, law enforcement, or advertising? Both Horikawa and Huth have stressed the importance of consent and privacy. \u201cNobody has shown you can do that, yet,\u201d Huth told Nature, when asked about reading private thoughts. But \u201cyet\u201d sounds concerning. <\/p>\n<p>For now, mind captioning is confined to the lab: a handful of subjects, a room-sized scanner, and a process that takes hours to calibrate. But the direction is unmistakable and hard to ignore. <\/p>\n<p>\t\t\t\t\t\t\t<script async src=\"https:\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n","protected":false},"excerpt":{"rendered":"Credit: Oryon. In a dark MRI scanner outside Tokyo, a volunteer watches a video of someone hurling themselves&hellip;\n","protected":false},"author":2,"featured_media":273738,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[256,254,255,64,63,158935,158936,158937,105],"class_list":["post-273737","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-au","tag-australia","tag-brain-scan","tag-mind-captioning","tag-mind-reading","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/273737","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=273737"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/273737\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/273738"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=273737"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=273737"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=273737"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}