{"id":847725,"date":"2026-09-14T17:52:16","date_gmt":"2026-09-14T17:52:16","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/847725\/"},"modified":"2026-09-14T17:52:16","modified_gmt":"2026-09-14T17:52:16","slug":"drug-firms-secret-data-supercharge-ai-protein-models","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/847725\/","title":{"rendered":"Drug firms\u2019 secret data supercharge AI protein models"},"content":{"rendered":"<p> <img decoding=\"async\" class=\"figure__image\" alt=\"Close-up of a protein ribbon structure, digital render wireframe with pink and blue helical structures on a black background.\" loading=\"lazy\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2026\/09\/d41586-026-02882-x_53696722.jpg\"\/><\/p>\n<p class=\"figure__caption u-sans-serif\">Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura\/Getty<\/p>\n<p>For drug discovery, protein-folding models such as AlphaFold have a data problem: there isn\u2019t enough of it in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.<\/p>\n<p>Protein structures\u2014 locked away by the thousands in drug company vaults\u2014 offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.<\/p>\n<p>The group used OpenFold3 \u2014 an open-source replication of AlphaFold 3 \u2014 to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available.<\/p>\n<p>\u201cYou add all this data, and you get a pretty big bump in performance,\u201d says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-025-00868-9\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2026\/09\/d41586-026-02882-x_52147804.png\"\/><\/p>\n<p class=\"recommended__title u-serif\">AlphaFold is running out of data \u2014 so drug firms are building their own version<\/p>\n<p><\/a><\/p>\n<p>The findings, he says, strengthen the case for generating similar publicly available datasets to supercharge protein-folding AIs. One such project, called OpenBind and supported by up to \u00a38 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month, with thousands more in the works.<\/p>\n<p>An untapped vein<\/p>\n<p>The Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, <a href=\"https:\/\/www.nature.com\/articles\/d41586-024-03423-0\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-024-03423-0\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">was the bedrock of AlphaFold 2\u2019s training data<\/a>. It enabled the tool to predict protein structures with startling accuracy \u2014 <a href=\"https:\/\/www.nature.com\/articles\/d41586-024-03214-7\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-024-03214-7\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">a breakthrough recognized<\/a><a href=\"https:\/\/www.nature.com\/articles\/d41586-024-03214-7\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-024-03214-7\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\"> with the 2024 Nobel Prize in Chemistry<\/a>.<\/p>\n<p>The model\u2019s successors, including AlphaFold 3, added <a href=\"https:\/\/www.nature.com\/articles\/d41586-024-01555-x\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-024-01555-x\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\">the ability to predict how proteins will interact with other molecules, including potential drugs<\/a>. But the PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules \u2014 maybe just 10,000, says Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.<\/p>\n<p>That lack of data is a problem for drug-discovery efforts. Research has suggested that the accuracy of AlphaFold 3 and other \u2018co-folding\u2019 models \u2014 which predict the structure of proteins interacting with each other \u2014 drops off a cliff when the models are challenged to predict interactions between molecules highly dissimilar to those on which they were trained<a href=\"#ref-CR1\" data-track=\"click\" data-action=\"anchor-link\" data-track-label=\"go to reference\" data-track-category=\"references\">1<\/a>.<\/p>\n<p>To make such tools more useful in drug discovery, researchers say, they need access to more data \u2014 which is why they have turned to vaults of molecular structures from pharmaceutical companies.<\/p>\n<p><a href=\"https:\/\/www.nature.com\/articles\/d41586-022-00997-5\" class=\"u-link-inherit\" data-track=\"select_article\" data-track-action=\"view recommended article\" data-track-context=\"related article widget on news body\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" class=\"recommended__image\" alt=\"\" src=\"https:\/\/www.newsbeep.com\/us\/wp-content\/uploads\/2026\/09\/d41586-026-02882-x_20323140.png\"\/><\/p>\n<p class=\"recommended__title u-serif\">What\u2019s next for AlphaFold and the AI protein-folding revolution<\/p>\n<p><\/a><\/p>\n<p>These data are generated during drug-discovery programmes, using techniques such as X-ray crystallography and cryo-electron microscopy. Many of the protein structures have never been deposited in public databases because they relate to proprietary drug-development efforts. The total size of these vaults is unknown, but some have estimated that they could contain more data than the PDB does.<\/p>\n<p>\u201cThe data that\u2019s missing from the PDB is exactly the data that\u2019s present in our internal data,\u201d John Karanicolas, head of computational drug discovery at the pharma company AbbVie in Chicago, Illinois, told Nature last year.<\/p>\n<p>Better predictions<\/p>\n<p>To test whether their data could be useful for protein-folding models, AbbVie, Astex and several other drug companies last year formed a collaboration<a href=\"https:\/\/www.nature.com\/articles\/d41586-025-00868-9\" data-track=\"click\" data-label=\"https:\/\/www.nature.com\/articles\/d41586-025-00868-9\" data-track-category=\"body text link\" rel=\"nofollow noopener\" target=\"_blank\"> called the AI Structural Biology (AISB) Network<\/a>.<\/p>\n<p>It involved \u2018fine-tuning\u2019 OpenFold3 \u2013 previously trained only with PDB data \u2013 on an additional 20,167 structures capturing proteins bound to potential drugs, or ligands. The structures came from five companies and were provided to the model in such a way that proprietary data remained private.<\/p>\n<p>The AISB study found that the extra data enhanced predictions. When tested on 1,056 protein\u2013ligand structures that were set aside from the training data, the AISB model predicted more than half of them to a high level of accuracy. By contrast, the publicly available version of OpenFold3 achieved the same performance on just one-third of the structures, and a competing open-source model called Boltz-2 achieved around 40%. The team plans to submit a paper describing the work to a peer-reviewed journal.<\/p>\n<p>The fact that the AISB model also outperformed co-folding tools that were trained only on each company\u2019s individual data highlights the benefits of pooling information, says Karanicolas.<\/p>\n","protected":false},"excerpt":{"rendered":"Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura\/Getty For drug&hellip;\n","protected":false},"author":2,"featured_media":847726,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[32],"tags":[3635,1159,1877,25214,1160,79],"class_list":["post-847725","post","type-post","status-publish","format-standard","has-post-thumbnail","category-science","tag-drug-discovery","tag-humanities-and-social-sciences","tag-machine-learning","tag-molecular-biology","tag-multidisciplinary","tag-science"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/847725","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=847725"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/847725\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/847726"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=847725"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=847725"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=847725"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}