{"id":459365,"date":"2026-09-10T12:50:14","date_gmt":"2026-09-10T12:50:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/us-ca\/459365\/"},"modified":"2026-09-10T12:50:14","modified_gmt":"2026-09-10T12:50:14","slug":"new-ai-model-for-dna-learns-from-evolution-to-study-the-human-genome","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us-ca\/459365\/","title":{"rendered":"New AI Model for DNA Learns from Evolution to Study the Human Genome"},"content":{"rendered":"<p class=\"p1 large\">UC Berkeley researchers have created a new genomic language model, called GPN-Star, that excels at spotting genetic variants that impact human health.<\/p>\n<p><img decoding=\"async\" alt=\"An illustration of a DNA strand.\" data-entity-type=\"file\" data-entity-uuid=\"5254eb2a-8dde-425a-b55a-0ec9ea2a93c3\" height=\"464\" src=\"https:\/\/www.newsbeep.com\/us-ca\/wp-content\/uploads\/2026\/09\/warren-umoh-KxwkcAe5Cpc-unsplash-768x444.jpeg\" width=\"803\" loading=\"lazy\"\/><br \/>\nLike chatbots for DNA, genomic language models are trained on vast troves of DNA sequences. Warren Umoh via. Unsplash+<\/p>\n<p class=\"p1\">More than two decades after scientists first sequenced the entire human genome \u2014\u00a0 all 3 billion \u201cletters,\u201d or base pairs, of DNA code \u2014 the\u00a0<a href=\"https:\/\/www.ucsf.edu\/news\/2017\/02\/405686\/mysterious-98-scientists-look-shine-light-our-dark-genome\" rel=\"nofollow noopener\" target=\"_blank\">meaning of much of this code<\/a>\u00a0remains a mystery.<\/p>\n<p class=\"p1\">While an estimated 1 to 2% of human DNA codes for proteins, the rest is a mix of \u201cjunk DNA\u201d \u2014 evolutionary holdovers that no longer code for anything \u2014 and regulatory elements that control when, where and how strongly genes are expressed. These non-coding regions of the genome could hold the key to understanding a variety of inherited traits, including those that lead to diseases such as cancer, heart disease and autism. But first, scientists have to understand how variants in this DNA contribute to the multitude of traits that make each of us unique.<\/p>\n<p class=\"p1\">Researchers at UC Berkeley have created\u00a0<a href=\"https:\/\/www.nature.com\/articles\/s41586-026-11005-5\" rel=\"nofollow noopener\" target=\"_blank\">a new genomic language AI model<\/a>, called GPN-Star, that far outpaces its competitors at identifying the most important genetic variants that contribute to inherited traits, including those that lead to disease. It is also far more computationally efficient than larger models, requiring only a fraction of the time and computing resources to train.<\/p>\n<p class=\"p1\">\u201cOur model excels in making predictions about the pathogenicity of genetic variants, and identifying functional versus non-functional elements in the genome,\u201d said study senior author\u00a0<a href=\"https:\/\/people.eecs.berkeley.edu\/~yss\/\" rel=\"nofollow noopener\" target=\"_blank\">Yun Song<\/a>, a professor of computer science and statistics at Berkeley and an investigator at the\u00a0<a href=\"https:\/\/innovativegenomics.org\/\" rel=\"nofollow noopener\" target=\"_blank\">Innovative Genomics Institute<\/a>.\u00a0<\/p>\n<p class=\"p1\">Along with the study, the researchers have published genome-wide predictions from their model, which highlight genetic variants that are likely to have the most influence on inherited traits. Biologists can use these annotations to identify relevant genes and regulatory elements for further study.<\/p>\n<p class=\"p1\">\u201cWe hope our work will help drive biological discovery,\u201d Song said. \u201cPeople have developed really creative tools for assaying the impact of genetic variants, but they cannot experimentally test every single variant in the genome. We believe our predictions will help to prioritize the experiments that could have the greatest impact on human health.\u201d<\/p>\n<p class=\"p1\">Song is also director of the\u00a0<a href=\"https:\/\/ccb.berkeley.edu\/\" rel=\"nofollow noopener\" target=\"_blank\">Berkeley Center for Computational Biology<\/a>\u00a0and co-director of the recently announced\u00a0<a href=\"https:\/\/ccb.berkeley.edu\/news\/yun-song-co-leads-uc-berkeley%E2%80%93ucsf-computational-biomedicine-initiative\" rel=\"nofollow noopener\" target=\"_blank\">UC Berkeley-UCSF Bakar Computational Biomedicine Initiative<\/a>. The study, funded in part by the National Institutes of Health, was published today (Sept. 9) in the journal\u00a0Nature.<\/p>\n<p><img decoding=\"async\" alt=\"Three people discuss a problem that is written on a white board.\" data-entity-type=\"file\" data-entity-uuid=\"89ae75a7-40dd-466d-8ef7-c44ad5f5beb9\" height=\"535\" src=\"https:\/\/www.newsbeep.com\/us-ca\/wp-content\/uploads\/2026\/09\/YunSongLab-768x512.jpg\" width=\"803\" loading=\"lazy\"\/><br \/>\nYun Song (left), Gonzolo Benegas and Chengzhong Ye in Song\u2019s office on the Berkeley campus. Glenn Ramit\/UC Berkeley<\/p>\n<p class=\"p2 large\">The grammar of the genome<\/p>\n<p class=\"p1\">Genomic language models work a little like chatbots for DNA, but instead of being trained on natural language, they are trained on vast troves of DNA sequences. These models\u2019 advanced pattern recognition skills can identify repeating patterns and sequences much faster than any human, allowing them to identify important elements of a genome that might otherwise be impossible to recognize.\u00a0<\/p>\n<p class=\"p1\">\u201cMathematically, a DNA sequence is just a string of letters \u2014 A, C, G and T. We don\u2019t know a priori which parts of the genome are functional elements, and a very small percentage of the genome is functional,\u201d Song said. \u201cBy training a DNA language model on a lot of different sequences, the model can recognize certain patterns that occur in the genomes. People have been using this to learn what we call the \u2018grammar\u2019 of the genome.\u201d<\/p>\n<p class=\"p1\">Most genomic language models, including the\u00a0<a href=\"https:\/\/www.nature.com\/articles\/d41586-025-00531-3\" rel=\"nofollow noopener\" target=\"_blank\">massive Evo 2 model<\/a>\u00a0published\u00a0<a href=\"https:\/\/www.nature.com\/articles\/s41586-026-10176-5\" rel=\"nofollow noopener\" target=\"_blank\">earlier this year<\/a>, are trained on sequences from entire unaligned genomes, which can range in size from humans all the way down to single-celled organisms. However, this approach can be extremely computationally demanding. The Evo 2 model, which can generate entire genomes from scratch, was trained on the genomes of more than 100,000 species across all domains of life, and required 2,000 powerful NVIDIA computer processors and months to train.\u00a0<\/p>\n<p class=\"p1\">To train the GPN-Star model, Song and his team used data from whole-genome alignments (WGAs) rather than individual unaligned genomes. WGAs use specialized algorithms to relate the genomes of hundreds of different species to that of a single species, highlighting similarities and differences in the code. For example, in a human-anchored WGA, the genomes of other species are compared to the human genome, revealing where the code has been conserved over the course of evolution and where it has changed.<\/p>\n<p class=\"p1\">Because WGAs do the work of identifying conserved areas of code, GPN-Star takes much less time and computing power to train than models that use unaligned genomes. It can be trained in just days, or even hours using only a handful of processors. It is also less likely to be confounded by the plethora of junk DNA that is found in the genomes of most species.\u00a0<\/p>\n<p class=\"p2 large\">Our approach is that we should use these biological insights to improve the model, rather than hoping that the model will figure out what\u2019s important by itself.<\/p>\n<p class=\"p1 large\">Professor Yun Song<\/p>\n<p class=\"p1\">\u201cWe tried to help the model learn by curating data that\u2019s more likely to harbor functional elements,\u201d Song said. \u201cOur approach is that we should use these biological insights to improve the model, rather than hoping that the model will figure out what\u2019s important by itself.\u201d\u00a0<\/p>\n<p class=\"p1\">In the new study, the team trained the model on three different human-anchored WGAs, as well as WGAs for mice, fruit flies, chickens,\u00a0C. elegans\u00a0(roundworms) and\u00a0A. thaliana\u00a0(a type of plant). Each of the human-anchored WGAs included genomes from a different combination of other species, representing different evolutionary timescales: One included the genomes of other primates, one included the genomes of mammals, and the final included the genomes of other vertebrates.<\/p>\n<p class=\"p1\">\u201cWe found that models trained at different evolutionary time scales were actually optimized for interpreting different kinds of genetic variants,\u201d said study co-first author\u00a0<a href=\"https:\/\/statistics.berkeley.edu\/people\/chengzhong-ye\" rel=\"nofollow noopener\" target=\"_blank\">Chengzhong Ye<\/a>, a graduate student in statistics at UC Berkeley. \u201cThis actually makes sense in terms of evolutionary biology, because some genomic elements evolve much faster than others.\u201d<\/p>\n<p class=\"p1\">For instance, they found that the model trained on data from longer evolutionary timescales was better at predicting the impact of rare genetic variants in proteins, which tend to evolve very slowly and are conserved across many species. However, the model trained on data from shorter evolutionary timescales was better at predicting the impact of genetic variants on complex traits like schizophrenia risk. These complex traits have been linked to as many as 10,000 different genetic mutations, many of which are in non-coding regions of the genome.\u00a0<\/p>\n<p class=\"p1\">\u201cFor complex traits, we were surprised and pleased to see that training a model that\u2019s specific to primate genomes \u2014 which are more relevant to recent human evolution \u2014 really helped us make better predictions,\u201d Song said.\u00a0<\/p>\n<p class=\"p1\">Because the model requires minimal resources to train, the researchers hope that it will be easy for other teams around the world to modify, adapt and improve upon their work, accelerating our understanding of genetics in humans and other species.\u00a0<\/p>\n<p class=\"p1\">\u201cWe\u2019re making great progress,\u201d said study co-first author\u00a0<a href=\"https:\/\/gonzalobenegas.github.io\/\" rel=\"nofollow noopener\" target=\"_blank\">Gonzalo Benegas<\/a>. \u201cBut the more people that can work with these models, the better they will get.\u201d<\/p>\n<p class=\"p1\">Additional co-authors of the study include Carlos Albors, Canal Li, Sebastian Prillo of Berkeley; Peter Fields of Jackson Laboratory; and Brian Clarke of the German Cancer Research Center in Heidelberg.\u00a0<\/p>\n<p class=\"p1\">This research was supported in part by National Institutes of Health grants R35-GM134922, R35- GM161566 and 3P40-OD011102-24S1 7772, and by the UC National Laboratory Fees Research Program of the University of California Office of the President (UC AI Science at Scale Grant L26CR10102). The Chan Zuckerberg Initiative provided GPU resources (through the \u201cAccelerating and Scaling Biological Sciences with AI\u201d program) to generate genome-wide predictions from the GPN-Star models.<\/p>\n","protected":false},"excerpt":{"rendered":"UC Berkeley researchers have created a new genomic language model, called GPN-Star, that excels at spotting genetic variants&hellip;\n","protected":false},"author":2,"featured_media":459366,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[139452,143,145,144,139450,139451],"class_list":["post-459365","post","type-post","status-publish","format-standard","has-post-thumbnail","category-oakland","tag-faculty-excellence","tag-oakland","tag-oakland-headlines","tag-oakland-news","tag-research-berkeley","tag-uc-berkeley-research"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/posts\/459365","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/comments?post=459365"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/posts\/459365\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/media\/459366"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/media?parent=459365"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/categories?post=459365"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us-ca\/wp-json\/wp\/v2\/tags?post=459365"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}