{"id":44142,"date":"2025-08-04T21:54:14","date_gmt":"2025-08-04T21:54:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/44142\/"},"modified":"2025-08-04T21:54:14","modified_gmt":"2025-08-04T21:54:14","slug":"deep-learning-based-gene-perturbation-effect-prediction-does-not-yet-outperform-simple-linear-baselines","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/44142\/","title":{"rendered":"Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines"},"content":{"rendered":"<p>The success of large language models in knowledge representation has spawned efforts to apply the foundation model concept to biology<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Gavriilidis, G. I., Vasileiou, V., Orfanou, A., Ishaque, N. &amp; Psomopoulos, F. A mini-review on perturbation modelling across single-cell omic modalities. Comput. Struct. Biotechnol. J. 23, 1886 (2024).\" href=\"#ref-CR1\" id=\"ref-link-section-d32254474e349\">1<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Sza&#x142;ata, A. et al. Transformers in single-cell omics: a review and new perspectives. Nat. Methods 21, 1430&#x2013;1443 (2024).\" href=\"#ref-CR2\" id=\"ref-link-section-d32254474e349_1\">2<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Rood, J. E., Hupalowska, A. &amp; Regev, A. Toward a foundation model of causal cell and tissue biology with a perturbation cell and tissue atlas. Cell 187, 4520&#x2013;4545 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR3\" id=\"ref-link-section-d32254474e352\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>. Several single-cell foundation models trained on transcriptomics data from millions of single cells have been published<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Yang, F. et al. scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data. Nat. Mach. Intell. 4, 852&#x2013;866 (2022).\" href=\"#ref-CR4\" id=\"ref-link-section-d32254474e356\">4<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Theodoris, C. V. et al. Transfer learning enables predictions in network biology. Nature 618, 616&#x2013;624 (2023).\" href=\"#ref-CR5\" id=\"ref-link-section-d32254474e356_1\">5<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Rosen, Y. et al. Universal cell embeddings: a foundation model for cell biology. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.11.28.568918&#010;                &#010;               (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR6\" id=\"ref-link-section-d32254474e359\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>. Two recent models\u2014scGPT<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Cui, H. et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nat. Methods 21, 1470&#x2013;1480 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR7\" id=\"ref-link-section-d32254474e363\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a> and scFoundation<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Hao, M. et al. Large-scale foundation model on single-cell transcriptomics. Nat. Methods 21, 1481&#x2013;1491 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR8\" id=\"ref-link-section-d32254474e367\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>\u2014claim to be able to predict gene expression changes caused by genetic perturbations.<\/p>\n<p>In the present study, we benchmarked the performance of these models against GEARS<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 9\" title=\"Roohani, Y., Huang, K. &amp; Leskovec, J. Predicting transcriptional outcomes of novel multigene perturbations with GEARS. Nat. Biotechnol. 42, 927&#x2013;935 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR9\" id=\"ref-link-section-d32254474e374\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a> and CPA<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Lotfollahi, M. et al. Predicting cellular responses to complex perturbations in high-throughput screens. Mol. Syst. Biol. 19, e11517 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR10\" id=\"ref-link-section-d32254474e378\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a> and against deliberately simplistic baselines. To provide additional perspective, we also included three single-cell foundation models\u2014scBERT<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Yang, F. et al. scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data. Nat. Mach. Intell. 4, 852&#x2013;866 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR4\" id=\"ref-link-section-d32254474e382\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>, Geneformer<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Theodoris, C. V. et al. Transfer learning enables predictions in network biology. Nature 618, 616&#x2013;624 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR5\" id=\"ref-link-section-d32254474e386\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> and UCE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Rosen, Y. et al. Universal cell embeddings: a foundation model for cell biology. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.11.28.568918&#010;                &#010;               (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR6\" id=\"ref-link-section-d32254474e390\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>\u2014that were not explicitly designed for this task but can be repurposed for it by combining them with a linear decoder that maps the cell embedding to the gene expression space. In the figures, we marked their results with an asterisk.<\/p>\n<p>We first assessed prediction of expression changes after double perturbations. We used data by Norman et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 11\" title=\"Norman, T. M. et al. Exploring genetic interaction manifolds constructed from rich single-cell phenotypes. Science 365, 786&#x2013;793 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR11\" id=\"ref-link-section-d32254474e397\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>, in which 100 individual genes and 124 pairs of genes were upregulated in K562 cells with a CRISPR activation system (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). The phenotypes for these 224 perturbations, plus the no-perturbation control, are logarithm-transformed RNA sequencing expression values for 19,264 genes.<\/p>\n<p>We fine-tuned the models on all 100 single perturbations and on 62 of the double perturbations and assessed the prediction error on the remaining 62 double perturbations. For robustness, we ran each analysis five times using different random partitions.<\/p>\n<p>For comparison, we included two simple baselines: (1) the \u2018no change\u2019 model that always predicts the same expression as in the control condition and (2) the \u2018additive\u2019 model that, for each double perturbation, predicts the sum of the individual logarithmic fold changes (LFCs). Neither uses the double perturbation data.<\/p>\n<p>All models had a prediction error substantially higher than the additive baseline (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1a,b<\/a>). Here, prediction error is the L2 distance between predicted and observed expression values for the 1,000 most highly expressed genes. We also examined other summary statistics, such as the Pearson delta measure, and L2 distances for other gene subsets: the n most highly expressed or the n most differentially expressed genes, for various n. We got the same overall result (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>).<\/p>\n<p>Fig. 1: Double perturbation prediction.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41592-025-02772-6\/figures\/1\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig1\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2025\/08\/41592_2025_2772_Fig1_HTML.png\" alt=\"figure 1\" loading=\"lazy\" width=\"685\" height=\"771\"\/><\/a><\/p>\n<p>a, Beeswarm plot of the prediction errors for 62 double perturbations across five test\u2013training splits. The prediction error is measured by the L2 distance between the predicted and the observed expression profile of the n\u2009=\u20091,000 most highly expressed genes. The horizontal red lines show the mean per model, which, for the best-performing model, is extended by the dashed line. b, Scatterplots of observed versus predicted expression from one example of the 62 double perturbations. The numbers indicate error measured by the L2 distance and the Pearson delta (R2). c, TPR (recall) of the interaction predictions as a function of the false discovery proportion. FN, false negative; FP, false positive; TP, true positive. d, Schematic of the classification of interactions based on the difference from the additive expectation (the error bars show the additive range). e, Bar chart of the composition of the observed interaction classes. f, Top: scatterplot of observed versus predicted expression compared to the additive expectation. Each point is one of the 1,000 read-out genes under one of the 62 double perturbations across five test\u2013training splits. The 500 predictions that deviated most from the additive expectation are depicted with bigger and more saturated points. Bottom: mosaic plots that compare the composition of highlighted predictions from the top panel stratified by the interaction class of the prediction. The width of the bars is scaled to match the number of instances. Source data for Fig. 1 are provided. expr., expression.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>Next, we considered the ability of the models to predict genetic interactions. Conceptually, a genetic interaction exists if the phenotype of two (or more) simultaneous perturbations is \u2018surprising\u2019. We operationalized this as double perturbation phenotypes that differed from the additive expectation more than expected under a null model with a Normal distribution (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Sec2\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). Using the full dataset, we identified 5,035 genetic interactions (out of potentially 124,000) at a false discovery rate of 5%.<\/p>\n<p>We then obtained genetic interaction predictions from each model by computing, for each of its 310,000 predictions (1,000 read-out genes and 62 held-out double perturbations across five test\u2013training splits), the difference between predicted expression and additive expectation, and, if that difference exceeded a given threshold D, we called a predicted interaction. We then computed, for all possible choices of D, the true-positive rate (TPR) and the false discovery proportion, which resulted in the curves shown in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1c<\/a>. The additive model did not compete as, by definition, it does not predict interactions.<\/p>\n<p>None of the models was better than the \u2018no change\u2019 baseline. The same ranking of models was observed when using other metrics (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>).<\/p>\n<p>To further dissect this finding, we classified the interactions as \u2018buffering\u2019, \u2018synergistic\u2019 or \u2018opposite\u2019 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1d,e<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Sec2\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). All models mostly predicted buffering interactions. The \u2018no change\u2019 baseline cannot, by definition, find synergistic interactions, but also the deep learning models rarely predicted synergistic interactions, and it was even rarer that those predictions were correct (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1f<\/a>).<\/p>\n<p>To our surprise, we often found the same pair of hemoglobin genes (HBG2 and HBZ) among the top predicted interactions, across models and double perturbations (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>). Examining the data, we noted that all models except Geneformer and scFoundation predicted LFC\u2009\u2248\u20090\u2014like the \u2018no change\u2019 baseline\u2014for the double perturbation of these two genes, despite their strong individual effects (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>). More generally, we noted that, for most genes, the predictions of scGPT, UCE and scBERT did not vary across perturbations, and those of GEARS and scFoundation varied considerably less than the ground truth (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>).<\/p>\n<p>GEARS, scGPT and scFoundation also claim the ability to predict the effect of unseen perturbations. GEARS uses shared Gene Ontology<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"The Gene Ontology Consortium. The Gene Ontology knowledgebase in 2023. Genetics 224, iyad031 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR12\" id=\"ref-link-section-d32254474e556\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a> annotations to extrapolate from the training data, whereas the foundation models are supposed to have learned the relationships between genes during pretraining to predict unseen perturbations.<\/p>\n<p>To benchmark this functionality, we used two CRISPR interference datasets by Replogle et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Replogle, J. M. et al. Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell 185, 2559&#x2013;2575 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR13\" id=\"ref-link-section-d32254474e563\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a> obtained with K562 and RPE1 cells and a dataset by Adamson et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Adamson, B. et al. A multiplexed single-cell CRISPR screening platform enables systematic dissection of the unfolded protein response. Cell 167, 1867&#x2013;1882 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR14\" id=\"ref-link-section-d32254474e567\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a> obtained with K562 cells (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>).<\/p>\n<p>As a baseline, we devised a simple linear model. It represents each read-out gene with a K-dimensional vector and each perturbation with an L-dimensional vector. These vectors are collected in the matrices G, with one row per read-out gene, and P, with one row per perturbation. G and P are either obtained as dimension-reducing embeddings of the training data (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Sec2\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>) or provided by an external source (see below). Then, given a data matrix Ytrain of gene expression values, with one row per read-out gene and one column per perturbation (that is, per condition pseudobulk of the single-cell data), the K\u2009\u00d7\u2009L matrix W is found as<\/p>\n<p>$$\\mathop{{\\rm{argmin}}}\\limits_{{\\bf{W}}}| | {{\\bf{Y}}}_{{\\rm{train}}}-({\\bf{G}}{\\bf{W}}{{\\bf{P}}}^{T}+{\\boldsymbol{b}})| {| }_{2}^{2},$$<\/p>\n<p>\n                    (1)\n                <\/p>\n<p>where b is the vector of row means of Ytrain (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>).<\/p>\n<p>Fig. 2: Single perturbation prediction.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41592-025-02772-6\/figures\/2\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig2\" src=\"https:\/\/www.newsbeep.com\/uk\/wp-content\/uploads\/2025\/08\/41592_2025_2772_Fig2_HTML.png\" alt=\"figure 2\" loading=\"lazy\" width=\"685\" height=\"720\"\/><\/a><\/p>\n<p>a, Beeswarm plot of the prediction errors for 134, 210 and 24 unseen single perturbations across two test\u2013training splits (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Sec2\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). The prediction error is measured by the L2 distance between the mean predicted and observed expression profile of the n\u2009=\u20091,000 most highly expressed genes. The horizontal red lines show the mean per model, which, for the best-performing model, is extended by the dashed line. DL, deep learning; LM, linear model. b, Schematic of the LM and how it can accommodate available gene (G) or perturbation (P) embeddings. c, Forest plot comparing the performance of all models relative to the error of the \u2018mean\u2019 baseline. The point ranges show the overall mean and 95% confidence interval of the bootstrapped mean ratio between each model and the baseline for 134, 210 and 24 unseen single perturbations across two test\u2013training splits. The opacity of the point range is reduced if the confidence interval contains 0. Source data for Fig. 2 are provided.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>We also included an even simpler baseline, b, the mean across the perturbations in the training set, following the preprints by Kernfeld et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Kernfeld, E., Yang, Y., Weinstock, J. S., Battle, A. &amp; Cahan, P. A systematic comparison of computational methods for expression forecasting. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.07.28.551039&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR15\" id=\"ref-link-section-d32254474e793\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a> and Csendes et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Csendes, G., Sanz, G., Szalay, K. Z. &amp; Szalai, B. Benchmarking foundation cell models for post-perturbation RNA-seq prediction. BMC Genomics 26, 393 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR16\" id=\"ref-link-section-d32254474e797\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a> that appeared while this paper was in revision.<\/p>\n<p>None of the deep learning models was able to consistently outperform the mean prediction or the linear model (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig10\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). We did not include scFoundation in this benchmark, as it required each dataset to exactly match the genes from its own pretraining data, and, for the Adamson and Replogle data, most of the required genes were missing. We also did not include CPA, as it is not designed to predict the effects of unseen perturbations.<\/p>\n<p>Next, we asked whether we could find utility in the data representations that GEARS, scGPT and scFoundation had learned during their pretraining. We extracted a gene embedding matrix G from scFoundation and scGPT, respectively, and a perturbation embedding matrix P from GEARS. The above linear model, equipped with these embeddings, performed as well or better than scGPT and GEARS with their in-built decoders (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2c<\/a>). Furthermore, the linear models with the gene embeddings from scFoundation and scGPT outperformed the \u2018mean\u2019 baseline, but they did not consistently outperform the linear model using G and P from the training data.<\/p>\n<p>The approach that did consistently outperform all other models was a linear model with P pretrained on the Replogle data (using the K562 cell line data as pretraining for the Adamson and RPE1 data and the RPE1 cell line for the K562 data). The predictions were more accurate for genes that were more similar between K562 and RPE1 (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>). Together, these results suggest that pretraining on the single-cell atlas data provided only a small benefit over random embeddings, but pretraining on perturbation data increased predictive performance.<\/p>\n<p>In summary, we presented prediction tasks where current foundation models did not perform better than deliberately simplistic linear prediction models, despite significant computational expenses for fine-tuning the deep learning models (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). As our deliberately simple baselines are incapable of representing realistic biological complexity, yet were not outperformed by the foundation models, we conclude that the latter\u2019s goal of providing a generalizable representation of cellular states and predicting the outcome of not-yet-performed experiments is still elusive.<\/p>\n<p>The publications that presented GEARS, scGPT and scFoundation included comparisons against GEARS and CPA and against a linear model. Some of these comparisons may have happened to be particularly \u2018easy\u2019. For instance, CPA was never designed to predict effects of unseen perturbations and was particularly uncompetitive in the double perturbation benchmark. The linear model used in scGPT\u2019s benchmark appears to have been set up such that it reverts to predicting no change over the control condition for any unseen perturbation.<\/p>\n<p>Our results are in line with previously published benchmarks that assessed the performance of foundation models for other tasks and found negligible benefits compared to simpler approaches<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Kedzierska, K. Z., Crawford, L., Amini, A. P. &amp; Lu, A. X. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biol. 26, 101 (2025).\" href=\"#ref-CR17\" id=\"ref-link-section-d32254474e850\">17<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Boiarsky, R., Singh, N., Buendia, A., Getz, G. &amp; Sontag, D. A deep dive into single-cell RNA sequencing foundation models. Preprint at bioRxiv &#10;                https:\/\/doi.org\/10.1101\/2023.10.19.563100&#10;                &#10;               (2023).\" href=\"#ref-CR18\" id=\"ref-link-section-d32254474e850_1\">18<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Liu, T., Li, K., Wang, Y., Li, H. &amp; Zhao, H. Evaluating the utilities of foundation models in single-cell data analysis. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.09.08.555192&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR19\" id=\"ref-link-section-d32254474e853\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>. Our results also concur with two previous studies showing that simple baselines outperform GEARS for predicting unseen single or double perturbations<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"M&#xE4;rtens, K., Donovan-Maiye, R. &amp; Ferkinghoff-Borg, J. Enhancing generative perturbation models with LLM-informed gene embeddings. In Proc. Workshop on Machine Learning for Genomics Explorations (ICLR, 2024). &#010;                https:\/\/openreview.net\/forum?id=eb3ndUlkt4&#010;                &#010;              \" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR20\" id=\"ref-link-section-d32254474e857\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Gaudelet, T. et al. Season combinatorial intervention predictions with salt &amp; peper. In Proc. Workshop on Machine Learning for Genomics Explorations (ICLR, 2024). &#010;                https:\/\/openreview.net\/forum?id=Wj95feICkN&#010;                &#010;              \" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR21\" id=\"ref-link-section-d32254474e860\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>. Since the release of our paper as a preprint, several other benchmarks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Kernfeld, E., Yang, Y., Weinstock, J. S., Battle, A. &amp; Cahan, P. A systematic comparison of computational methods for expression forecasting. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.07.28.551039&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR15\" id=\"ref-link-section-d32254474e864\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Csendes, G., Sanz, G., Szalay, K. Z. &amp; Szalai, B. Benchmarking foundation cell models for post-perturbation RNA-seq prediction. BMC Genomics 26, 393 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR16\" id=\"ref-link-section-d32254474e867\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wenteler, A. et al. PertEval-scFM: benchmarking single-cell foundation models for perturbation effect prediction. In Proc. 42nd International Conference on Machine Learning (ICML, 2025).\" href=\"#ref-CR22\" id=\"ref-link-section-d32254474e870\">22<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Bendidi, I. et al. Benchmarking transcriptomics foundation models for perturbation analysis: one PCA still rules them all. Preprint at &#10;                https:\/\/arxiv.org\/abs\/2410.13956&#10;                &#10;               (2024).\" href=\"#ref-CR23\" id=\"ref-link-section-d32254474e870_1\">23<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wu, Y. et al. PerturBench: benchmarking machine learning models for cellular perturbation analysis. Preprint at &#10;                https:\/\/arxiv.org\/abs\/2408.10609&#10;                &#10;               (2024).\" href=\"#ref-CR24\" id=\"ref-link-section-d32254474e870_2\">24<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Li, L. et al. A systematic comparison of single-cell perturbation response prediction models. Preprint at bioRxiv &#10;                https:\/\/doi.org\/10.1101\/2024.12.23.630036&#10;                &#10;               (2024).\" href=\"#ref-CR25\" id=\"ref-link-section-d32254474e870_3\">25<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Li, C. et al. Benchmarking AI models for in silico gene perturbation of cells. Preprint at bioRxiv &#10;                https:\/\/doi.org\/10.1101\/2024.12.20.629581&#10;                &#10;               (2025).\" href=\"#ref-CR26\" id=\"ref-link-section-d32254474e870_4\">26<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Wong, D. R., Hill, A. S. &amp; Moccia, R. Simple controls exceed best deep learning algorithms and reveal foundation model effectiveness for predicting genetic perturbations. Bioinformatics 41, btaf317 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR27\" id=\"ref-link-section-d32254474e873\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a> were released that also show that deep learning models struggle to outperform simple baselines. Two of these preprints<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Kernfeld, E., Yang, Y., Weinstock, J. S., Battle, A. &amp; Cahan, P. A systematic comparison of computational methods for expression forecasting. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2023.07.28.551039&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR15\" id=\"ref-link-section-d32254474e877\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Csendes, G., Sanz, G., Szalay, K. Z. &amp; Szalai, B. Benchmarking foundation cell models for post-perturbation RNA-seq prediction. BMC Genomics 26, 393 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR16\" id=\"ref-link-section-d32254474e880\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a> suggested an even simpler baseline than our linear model (equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#Equ1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>)), namely, to always predict the overall average, and we have included this idea here.<\/p>\n<p>One limitation of our benchmark is that we used only four datasets. We chose these as they were used in the publications presenting GEARS, scGPT and scFoundation. Another limitation is that all datasets are from cancer cell lines, which, for example, Theodoris et al.<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Theodoris, C. V. et al. Transfer learning enables predictions in network biology. Nature 618, 616&#x2013;624 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR5\" id=\"ref-link-section-d32254474e890\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> excluded from their training data because of concerns about their high mutational burden. We also did not attempt to improve the original quality control, for example, by excluding perturbations that did not affect the expression of their own target gene and, thus, might not have worked as intended.<\/p>\n<p>Deep learning is effective in many areas of single-cell omics<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Gayoso, A. et al. A Python library for probabilistic analysis of single-cell omics data. Nat. Biotechnol. 40, 163&#x2013;166 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR28\" id=\"ref-link-section-d32254474e898\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Luecken, M. D. et al. Benchmarking atlas-level data integration in single-cell genomics. Nat. Methods 19, 41&#x2013;50 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41592-025-02772-6#ref-CR29\" id=\"ref-link-section-d32254474e901\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>. However, prediction of perturbation effects still remains an open challenge, as our present work shows. We expect that increased focus on performance metrics and benchmarking will be instrumental to facilitate eventual success in applying transfer learning to perturbation data.<\/p>\n","protected":false},"excerpt":{"rendered":"The success of large language models in knowledge representation has spawned efforts to apply the foundation model concept&hellip;\n","protected":false},"author":2,"featured_media":44143,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,10401,23002,23001,18924,3250,6194,11376,5395,4863,86,11804,56,54,55],"class_list":["post-44142","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-bioinformatics","tag-biological-microscopy","tag-biological-techniques","tag-biomedical-engineering-biotechnology","tag-general","tag-life-sciences","tag-machine-learning","tag-proteomics","tag-software","tag-technology","tag-transcriptomics","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/44142","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=44142"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/44142\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/44143"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=44142"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=44142"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=44142"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}