{"id":627882,"date":"2026-05-07T00:59:11","date_gmt":"2026-05-07T00:59:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/627882\/"},"modified":"2026-05-07T00:59:11","modified_gmt":"2026-05-07T00:59:11","slug":"non-invasive-profiling-of-the-tumour-microenvironment-with-spatial-ecotypes","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/627882\/","title":{"rendered":"Non-invasive profiling of the tumour microenvironment with spatial ecotypes"},"content":{"rendered":"<p>Human subjects<\/p>\n<p>All human samples included in this study were collected with informed consent for research use and received approval from the Institutional Review Boards of Yale University School of Medicine and Washington University School of Medicine, in accordance with the principles of the Declaration of Helsinki (2013). These samples, obtained from a total of 123 human subjects, were divided into four cohorts.<\/p>\n<p>Cohort 1<\/p>\n<p>Intact and dissociated tumour samples were collected from seven patients (four with colon cancer and three with melanoma) at the time of surgery. Each sample underwent bulk RNA sequencing and the dissociated tumour samples also underwent scRNA-seq.<\/p>\n<p>Cohort 2<\/p>\n<p>Matched tumour and plasma cfDNA samples were collected from 23 patients with\u00a0metastatic melanoma, with matched PBMCs also collected for seven. Tumour samples for each patient were profiled by ST and\/or whole-genome EM-seq, depending on availability (Visium, Visium HD and\/or EM-seq). A further 23 plasma cfDNA samples were collected from healthy individuals. All PBMC and plasma cfDNA samples were profiled by whole-genome EM-seq. Matched melanoma samples are graphically depicted in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4d<\/a> and a full inventory is provided in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>.<\/p>\n<p>Cohort 3<\/p>\n<p>Plasma samples were collected from 78 patients with melanoma\u00a0treated at Yale Cancer Center, including seven from cohort 2, who received ICI monotherapy (30 received anti-PD-1 and five received anti-CTLA-4) or combination therapy (43 received anti-PD-1 and anti-CTLA-4). Samples were collected before treatment initiation (before or on the first day of ICI cycle 1) and underwent whole-genome EM-seq. ICI response was classified as either durable clinical benefit or no durable benefit by a board-certified medical oncologist, reflecting each patient\u2019s disease response six months after ICI initiation. Progression-free survival was determined from the start of ICI treatment.<\/p>\n<p>Cohort 4<\/p>\n<p>Plasma samples were collected from ten patients with melanoma\u00a0treated at Siteman Cancer Center, including eight from cohort 2, who received immune checkpoint inhibitor (ICI) monotherapy (seven received anti-PD-1) or combination therapy (three received anti-PD-1 and anti-CTLA-4). Samples were collected before or during treatment and underwent whole-genome EM-seq. ICI response was classified as described for cohort 3.<\/p>\n<p>All clinical features, including age and sex, were documented using electronic medical records from Siteman Cancer Center (cohorts 1, 2 and 4) and Yale Cancer Center (cohorts 2 and 3). De-identified clinical characteristics are provided in Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a> for the four cohorts. Details of sample processing and sequencing are provided in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>Data collection and processingscRNA-seq<\/p>\n<p>Single-cell RNA-seq atlases of carcinomas and melanomas, either generated in this work or obtained from published studies as preprocessed data<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Tirosh, I. et al. Dissecting the multicellular ecosystem of metastatic melanoma by single-cell RNA-seq. Science 352, 189&#x2013;196 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR4\" id=\"ref-link-section-d74592869e3206\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Qi, J. et al. Single-cell and spatial analysis reveal interaction of FAP+ fibroblasts and SPP1+ macrophages in colorectal cancer. Nat. Commun. 13, 1742 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR27\" id=\"ref-link-section-d74592869e3209\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Xing, X. et al. Pan-cancer human brain metastases atlas at single-cell resolution. Cancer Cell 43, 1242&#x2013;1260 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR36\" id=\"ref-link-section-d74592869e3212\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Chen, S. et al. Single-cell analysis reveals transcriptomic remodellings in distinct cell types that contribute to human prostate cancer progression. Nat. Cell Biol. 23, 87&#x2013;98 (2021).\" href=\"#ref-CR59\" id=\"ref-link-section-d74592869e3215\">59<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wu, S. Z. et al. A single-cell and spatially resolved atlas of human breast cancers. Nat. Genet. 53, 1334&#x2013;1347 (2021).\" href=\"#ref-CR60\" id=\"ref-link-section-d74592869e3215_1\">60<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wu, F. et al. Single-cell profiling of tumor heterogeneity and the microenvironment in advanced non-small cell lung cancer. Nat. Commun. 12, 2540 (2021).\" href=\"#ref-CR61\" id=\"ref-link-section-d74592869e3215_2\">61<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Lu, Y. et al. A single-cell atlas of the multicellular ecosystem of primary and metastatic hepatocellular carcinoma. Nat. Commun. 13, 4594 (2022).\" href=\"#ref-CR62\" id=\"ref-link-section-d74592869e3215_3\">62<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Olalekan, S., Xie, B., Back, R., Eckart, H. &amp; Basu, A. Characterizing the tumor microenvironment of metastatic ovarian cancer by single-cell transcriptomics. Cell Rep. 35, 109165 (2021).\" href=\"#ref-CR63\" id=\"ref-link-section-d74592869e3215_4\">63<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Peng, J. et al. Single-cell RNA-seq highlights intra-tumoral heterogeneity and malignant progression in pancreatic ductal adenocarcinoma. Cell Res. 29, 725&#x2013;738 (2019).\" href=\"#ref-CR64\" id=\"ref-link-section-d74592869e3215_5\">64<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Lai, H. et al. Single-cell RNA sequencing reveals the epithelial cell heterogeneity and invasive subpopulation in human bladder cancer. Int. J. Cancer 149, 2099&#x2013;2115 (2021).\" href=\"#ref-CR65\" id=\"ref-link-section-d74592869e3215_6\">65<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 66\" title=\"Ji, A. L. et al. Multimodal analysis of composition and spatial architecture in human squamous cell carcinoma. Cell 182, 497&#x2013;514 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR66\" id=\"ref-link-section-d74592869e3218\" rel=\"nofollow noopener\" target=\"_blank\">66<\/a> (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>), were annotated for B cells, plasma cells, CD4+ T cells, CD8+ T cells, NK cells, macrophages\u00a0(including monocytes), dendritic cells, fibroblasts, endothelial cells, epithelial or melanoma cells and other\u00a0(unclassified)\u00a0cell types. For publicly available datasets with author-supplied annotations (breast cancer, colon cancer, liver cancer, squamous cell carcinoma, melanoma and brain metastasis), annotations were mapped to the above cell-type labels. Cohort 1 scRNA-seq data, as well as publicly available data without cell-type annotations (bladder cancer, lung cancer, ovarian cancer, prostate cancer and pancreatic cancer), were analysed using Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e3233\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a> as described below. For quality control, cells with fewer than 200 detected genes or more than 25% of reads mapped to mitochondrial genes were excluded. Raw counts were imported and cells clustered following SCTransform with the glmGamPoi method, FindVariableFeatures, ScaleData, RunPCA, FindNeighbors and FindClusters. Cell-type annotations were then manually assigned to clusters based on the expression of canonical lineage markers (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1c<\/a>): MS4A1, CD19, CD79A and CD79B for B cells, IGKC and MZB1 for plasma cells, CD3D, CD4 and IL7R for CD4+ T cells, CD3D, CD8A and CD8B for CD8+ T cells, GNLY and NCAM1 for NK cells, CD68 and CD14 for macrophages\/monocytes, CD1C for dendritic cells, COL1A1, COL3A1, PDGFRA and FAP for fibroblasts, PECAM1 and VWF for endothelial cells, EPCAM for epithelial cells and SOX9, MET, MITF and MLANA for melanoma cells. Small clusters with multilineage marker expression were considered potential doublets or multiplets and eliminated from further analysis (3% of filtered cells from cohort 1, on average).<\/p>\n<p>Visium and legacy ST<\/p>\n<p>Processed data from 54 Visium (standard) and 54 legacy ST profiles of carcinoma and melanoma samples were downloaded from 10x Genomics (<a href=\"https:\/\/www.10xgenomics.com\/resources\/datasets\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.10xgenomics.com\/resources\/datasets<\/a>) and 12 previous studies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Qi, J. et al. Single-cell and spatial analysis reveal interaction of FAP+ fibroblasts and SPP1+ macrophages in colorectal cancer. Nat. Commun. 13, 1742 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR27\" id=\"ref-link-section-d74592869e3348\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 60\" title=\"Wu, S. Z. et al. A single-cell and spatially resolved atlas of human breast cancers. Nat. Genet. 53, 1334&#x2013;1347 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR60\" id=\"ref-link-section-d74592869e3351\" rel=\"nofollow noopener\" target=\"_blank\">60<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 66\" title=\"Ji, A. L. et al. Multimodal analysis of composition and spatial architecture in human squamous cell carcinoma. Cell 182, 497&#x2013;514 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR66\" id=\"ref-link-section-d74592869e3354\" rel=\"nofollow noopener\" target=\"_blank\">66<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Gouin, K. H. et al. An N-Cadherin 2 expressing epithelial cell subpopulation predicts response to surgery, chemotherapy and immunotherapy in bladder cancer. Nat. Commun. 12, 4906 (2021).\" href=\"#ref-CR68\" id=\"ref-link-section-d74592869e3357\">68<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Wu, R. et al. Comprehensive analysis of spatial architecture in primary liver cancer. Sci. Adv. 7, eabg3750 (2021).\" href=\"#ref-CR69\" id=\"ref-link-section-d74592869e3357_1\">69<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Barkley, D. et al. Cancer cell states recur across tumor types and form specific interactions with the tumor microenvironment. Nat. Genet. 54, 1192&#x2013;1201 (2022).\" href=\"#ref-CR70\" id=\"ref-link-section-d74592869e3357_2\">70<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Thrane, K., Eriksson, H., Maaskola, J., Hansson, J. &amp; Lundeberg, J. Spatially resolved transcriptomics enables dissection of genetic heterogeneity in stage III cutaneous malignant melanoma. Cancer Res. 78, 5970&#x2013;5979 (2018).\" href=\"#ref-CR71\" id=\"ref-link-section-d74592869e3357_3\">71<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Moncada, R. et al. Integrating microarray-based spatial transcriptomics and single-cell RNA-seq reveals tissue architecture in pancreatic ductal adenocarcinomas. Nat. Biotechnol. 38, 333&#x2013;342 (2020).\" href=\"#ref-CR72\" id=\"ref-link-section-d74592869e3357_4\">72<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Berglund, E. et al. Spatial maps of prostate cancer transcriptomes reveal an unexplored landscape of heterogeneity. Nat. Commun. 9, 2419 (2018).\" href=\"#ref-CR73\" id=\"ref-link-section-d74592869e3357_5\">73<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Berglund, E. et al. Automation of Spatial Transcriptomics library preparation to enable rapid and robust insights into spatial organization of tissues. BMC Genomics 21, 298 (2020).\" href=\"#ref-CR74\" id=\"ref-link-section-d74592869e3357_6\">74<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Andersson, A. et al. Spatial deconvolution of HER2-positive breast cancer delineates tumor-associated cell type interactions. Nat. Commun. 12, 6012 (2021).\" href=\"#ref-CR75\" id=\"ref-link-section-d74592869e3357_7\">75<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 76\" title=\"He, B. et al. Integrating spatial gene expression and breast tumour morphology via deep learning. Nat. Biomed. Eng. 4, 827&#x2013;834 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR76\" id=\"ref-link-section-d74592869e3360\" rel=\"nofollow noopener\" target=\"_blank\">76<\/a> (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>). Legacy ST refers to the predecessor of 10x Visium, a lower-resolution ST assay with 100\u2009\u03bcm spot diameter reported in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 77\" title=\"St&#xE5;hl, P. L. et al. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science 353, 78&#x2013;82 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR77\" id=\"ref-link-section-d74592869e3370\" rel=\"nofollow noopener\" target=\"_blank\">77<\/a>. For quality control of publicly available data, genes expressed in fewer than five spots and spots expressing fewer than 200 unique genes were omitted. Processing details of Visium data generated in this study are provided in \u201810x Visium (standard)\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>MERSCOPE<\/p>\n<p>Preprocessed MERSCOPE profiles of 15 FFPE human tumour specimens, spanning melanoma and six distinct carcinomas, were downloaded from Vizgen (MERSCOPE FFPE Human Immuno-oncology program; <a href=\"https:\/\/info.vizgen.com\/merscope-ffpe-solution\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/info.vizgen.com\/merscope-ffpe-solution<\/a>). Three ovarian cancer samples were excluded owing to substantial tissue fragmentation (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). For quality control, genes expressed in fewer than five cells and cells with fewer than 300 total transcripts were excluded from each remaining sample.<\/p>\n<p>For cell-type annotation, transcripts were downsampled to 300 per cell, and cells were clustered with Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e3399\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a> using the following steps: NormalizeData, FindVariableFeatures (nfeatures\u2009=\u2009300), ScaleData, RunPCA, FindNeighbors and FindClusters (resolution\u2009=\u20091). Cell-type annotations were then assigned by cluster based on the expression of canonical markers (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a>) as described for scRNA-seq above, with reclustering of individual clusters, particularly those containing mixed lymphocyte groups, performed for greater granularity as needed.<\/p>\n<p>To remove ambient and improperly segmented mRNAs, publicly available tumour scRNA-seq atlases (\u2018scRNA-seq\u2019 above) were used to identify genes commonly expressed in each cell type. Specifically, for each cell type, genes expressed in at least 5% of cells in three or more cancer types were identified, resulting in a whitelist for each cell type. Genes in each cell type that were absent from the corresponding whitelist were then set to zero expression in the MERSCOPE data. Subsequently, cells with fewer than five detectably expressed genes were excluded, resulting in a final dataset of 5.6\u2009million evaluable cells from 12 samples (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). For MERSCOPE data generated in this study, see \u2018Vizgen MERSCOPE\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>Xenium<\/p>\n<p>Space Ranger results from 11 samples (nine carcinomas and two melanomas) profiled with Xenium V1 (n\u2009=\u20095) and Xenium Prime (n\u2009=\u20096) were downloaded from <a href=\"https:\/\/www.10xgenomics.com\/resources\/datasets\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.10xgenomics.com\/resources\/datasets<\/a> (Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). Cell-type annotation by canonical marker expression within clusters (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a>) and subsequent postprocessing were done following the workflow described in \u2018MERSCOPE\u2019\u00a0above, with two modifications, owing to the lower average number of detected genes in Xenium compared with MERSCOPE: omission of the downsampling step to 300 transcripts per cell; and variable feature selection using FindVariableFeatures with nfeatures\u2009=\u2009200. One sample lacking annotatable CD4+ T cells was excluded from downstream analysis (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). For quality control, cells with fewer than 20 detected genes or with greater than 10% mitochondrial transcript content were omitted.<\/p>\n<p>Visium HD<\/p>\n<p>Space Ranger results from five carcinoma samples profiled with Visium HD (bins of 8\u2009\u03bcm\u2009\u00d7\u20098\u2009\u03bcm) were downloaded from <a href=\"https:\/\/www.10xgenomics.com\/resources\/datasets\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.10xgenomics.com\/resources\/datasets<\/a> (Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). Cell-type annotation and postprocessing were performed as described for Xenium data (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a>). For these samples, Visium HD bins generally yielded robust cell-type discrimination, in line with a previous report<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"Oliveira, M. F. de et al. High-definition spatial transcriptomic profiling of immune cell populations in colorectal cancer. Nat. Genet. 57, 1512&#x2013;1523 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR78\" id=\"ref-link-section-d74592869e3482\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a>. One sample lacking annotatable CD8+ T cells was excluded from downstream analysis (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). For quality control of publicly available data, bins with fewer than 20 detected genes or with greater than 10% mitochondrial transcript content were omitted. Processing details for Visium HD data generated in this study, which were used for SE deconvolution as described below in section\u00a0\u2018Paired tumour and plasma from patients with\u00a0melanoma\u2019, are provided in \u201810x Visium HD\u2019 in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>Analysis of tumour versus adjacent stromaIntegration of single-cell and spatial transcriptomes<\/p>\n<p>CytoSPACE (v.1.0.3)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Vahid, M. R. et al. High-resolution alignment of single-cell and spatial transcriptomes with CytoSPACE. Nat. Biotechnol. 41, 1543&#x2013;1548 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR25\" id=\"ref-link-section-d74592869e3508\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a> was used to align scRNA-seq data to ST data from the same cancer type, reconstructing transcriptome-wide spatially resolved expression profiles of single cells (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1c<\/a>). Alignment was done separately for each ST sample, with source data enumerated in Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. To eliminate potential bias arising from different total unique molecular identifiers (UMIs), raw counts from droplet-based scRNA-seq data were downsampled to 1,500 total UMIs per cell, whereas transcripts per million (TPM) data from Smart-seq2 samples (melanoma) were used without downsampling. The lap_CSPR solver and recommended settings were applied for all analyses, including the default mode for bulk ST (Visium and legacy ST), single-cell mode for single-cell ST data (MERSCOPE) and an average of five cells per spot for Visium data and 20 cells per spot for legacy ST data.<\/p>\n<p>Differential expression analysis<\/p>\n<p>For the results presented in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1d<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1a\u2013d,e,g<\/a>, we analysed scRNA-seq data mapped to ST samples, as described above. We also analysed MERSCOPE data directly (without scRNA-seq integration) as a form of reciprocal validation (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1a\u2013c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). As input to the analyses described below, UMI-based scRNA-seq data were normalized for each cell type separately using SCTransform from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e3541\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>; Smart-seq2 data from melanoma samples were normalized to log2[TPM]; and MERSCOPE data were normalized using NormalizeData from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e3548\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>.<\/p>\n<p>To study transcriptome-wide variation in TME cell types localized to the tumour or adjacent stroma, and given broad cancer coverage and sample availability, CytoSPACE-enhanced Visium data were selected as the primary discovery cohort. Differential expression between tumour and adjacent stroma (annotated as described in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>) was first determined for each cell type and sample separately using the wilcoxauc function from presto (v.1.1.0; <a href=\"https:\/\/github.com\/immunogenomics\/presto\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/immunogenomics\/presto<\/a>)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"Korsunsky, I., Nathan, A., Millard, N. &amp; Raychaudhuri, S. Presto scales Wilcoxon and auROC analyses to millions of observations. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/653253&#010;                &#010;               (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR79\" id=\"ref-link-section-d74592869e3565\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>. Log2-transformed fold changes (LFCs) for each gene were then aggregated by median to avoid bias, and corresponding meta P-values were calculated using Stouffer\u2019s approach<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Stouffer, S. A., Suchman, E. A., Devinney, L. C., Star, S. A. &amp; Williams, R. M. Jr. The American Soldier: Adjustment during Army Life. (Studies in Social Psychology in World War II) Vol. 1 (Princeton Univ. Press, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR80\" id=\"ref-link-section-d74592869e3575\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a> following conversion of two-sided P-values to z-scores. The calculations were done first across sample replicates, then across all samples within each cancer type, and finally across all cancer types in the discovery cohort. Meta P-values were adjusted using the Benjamini\u2013Hochberg method to derive Q-values<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 81\" title=\"Benjamini, Y. &amp; Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. B Methodol. 57, 289&#x2013;300 (1995).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR81\" id=\"ref-link-section-d74592869e3591\" rel=\"nofollow noopener\" target=\"_blank\">81<\/a>. Significantly differentially expressed genes between tumour and adjacent stroma were identified as genes with: significant differential expression (per-cancer LFC\u2009&gt;\u20090.05 and Q\u2009&lt;\u20090.05) in at least three cancer types; a pan-cancer Q\u2009&lt;\u20090.05; and a pan-cancer median LFC\u2009&gt;\u20090.02. Genes with conserved pan-cell-type enrichment were omitted from this analysis and examined elsewhere (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1e<\/a>; see below). Among the remaining genes, up to 400 HUGO protein-coding genes (<a href=\"https:\/\/www.genenames.org\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.genenames.org<\/a>) with the highest LFC in tumour (n\u2009\\(\\le \\)\u2009200) or adjacent stroma (n\u2009\\(\\le \\)\u2009200) were visualized across all held-out samples mapped to scRNA-seq data (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1d<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1d<\/a>). For cell types with fewer than 200 genes in either region, the minimum number per compartment was selected for balance (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). For visualization, data were scaled per column to a maximum absolute LFC of one and genes were ordered by the resulting median enrichment balanced across platforms (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1d<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1d<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>).<\/p>\n<p>To cross-validate CytoSPACE-enhanced Visium against single-cell ST data directly (without scRNA-seq integration), we repeated the above analysis using the 500 genes covered by the MERSCOPE panel (section\u00a0\u2018MERSCOPE\u2019\u00a0above). We then identified up to 50 HUGO protein-coding genes by median LFC (tumour, n\u2009\\(\\le \\)\u200925; adjacent stroma, n\u2009\\(\\le \\)\u200925) for each TME cell type and repeated this step independently for each platform. Cross-platform concordance was quantified by Spearman correlation of median LFCs and by the directionality of expression (higher in tumour or adjacent stroma) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1a\u2013c<\/a>). All results are detailed in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>.<\/p>\n<p>To identify spatially polarized genes with conservation across TME cell types, genes with differential expression (pan-cancer LFC\u2009&gt;\u20090.05 and Q\u2009&lt;\u20090.05) in more than 50% of cell types (n\u2009=\u20095) in the Visium discovery cohort were ranked by average LFC across all cell types. For balanced representation, the minimum number of top-ranking genes per compartment was selected for visualization (tumour, n\u2009=\u2009138; adjacent stroma, n\u2009=\u2009138) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">1e<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>).<\/p>\n<p>Spatial EcoTyper framework<\/p>\n<p>Despite experimental advances enabling high-resolution expression profiling of cells in situ, leveraging such data to systematically profile the co-association of cell states into SEs and discover conserved SEs across specimens and cancer types has remained challenging. The spatial organization of cell states and their relative abundances in an ecotype can vary across regions and between samples, and even expression profiles of individual cells sharing the same phenotypic state exhibit natural variability. Furthermore, technical drop-out and sample- or platform-specific batch effects all pose obstacles to SE discovery.<\/p>\n<p>With these considerations in mind, we developed Spatial EcoTyper (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). At its core, the framework relies upon a network integration technique to identify common patterns of ST variation shared across samples. This is achieved by adapting similarity network fusion (SNF), a previously described approach for multi-omics data integration across patients<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Wang, B. et al. Similarity network fusion for aggregating data types on a genomic scale. Nat. Methods 11, 333&#x2013;337 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR32\" id=\"ref-link-section-d74592869e3741\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>. By introducing a series of carefully constructed spatial GEPs, our approach mitigates technical drop-out while providing stability under biological variation. Once defined, SEs can be robustly recovered in a supervised manner from non-spatial data using unique cell states and molecular signatures that are learnt from spatial data.<\/p>\n<p>Spatial EcoTyper consists of five key components, described in detail in the following sections.<\/p>\n<p>Determination of sample-level spatial clusters. In each single-cell ST sample, spatial expression data are encoded into cell-type-specific GEPs of spatial neighbourhoods (SNs), spatial covariation among the SNs is computed by SNF and SNs are clustered over the resulting network (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>; steps 1\u20134, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>).<\/p>\n<p>Identification of conserved SEs. Spatial clusters discovered from individual single-cell ST samples are represented by GEPs of their associated cell states, and clusters with similar GEPs are aggregated across samples into conserved SEs (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>; steps 1\u20139, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>).<\/p>\n<p>Discovery of conserved SE-specific cell states. Cell states uniquely enriched in each SE and conserved across samples are identified using a specialized variant of NMF (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6a<\/a>).<\/p>\n<p>Recovery of conserved SE-specific cell states. An NMF model is developed to recover SE-specific cell states in external single-cell or spatial expression datasets (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6b<\/a>).<\/p>\n<p>Deconvolution of SEs from bulk RNA-seq. The approach from the previous component is generalized to the task of recovering SE abundances from bulk RNA-seq, using a training cohort of pseudo-bulk mixtures (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>).<\/p>\n<p>              Determination of sample-level spatial clusters<\/p>\n<p>The Spatial EcoTyper framework, schematically illustrated in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>, begins by identifying clusters of SNs in each single-cell ST sample (steps 1\u20133, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). Although ST data should be generated by the same assay, the discovery phase is applicable to diverse single-cell ST platforms. In this work, we used tumour samples profiled by MERSCOPE for discovery (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>), with normalization performed as described in the \u2018Spatial EcoTyper discovery cohort\u2019\u00a0section\u00a0below. To assess reproducibility, we also applied discovery mode to Xenium Prime data as described in \u2018Robustness of spatial ecotype discovery to single-cell ST platform\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a> (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6k\u2013n<\/a>).<\/p>\n<p>Assembly of cell-type-specific SN expression profiles<\/p>\n<p>Spatial proximity is explicitly used by Spatial EcoTyper in two ways: when analysing individual cells, and when analysing distinct cell types. To accomplish the former, cell-type-specific GEPs are first aggregated by SNs centred along a regular grid (step 1, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). This involves constructing a vector for each cell type c, denoted snGEPc, by averaging the normalized GEPs of the nearest up to k cells of cell type c located within radius r of the centre of each SN (selected to be 50\u2009\u03bcm in practice; section \u2018SN radius\u2019 below). The snGEPc vectors for all SNs are then concatenated into matrix Ec with g genes (rows) and m snGEPc vectors (columns). Crucially, the latter is consistently ordered left-to-right by SN coordinates, enabling co-registration across cell types. In this way, for any given SN with coordinates i, j, snGEP vectors for cell types x and y will occupy the same column index in Ex and Ey, respectively (step 1, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). To identify SEs containing multiple cell types, SNs characterized by a single cell type are eliminated by default. Furthermore, from each Ec matrix, genes expressed in fewer than five SNs, SNs with no cells of type c and SNs expressing fewer than five genes are excluded with associated entries set to NA.<\/p>\n<p>The snGEPc vector serves as a fundamental data unit for Spatial EcoTyper, analogous to a \u2018spatial meta-cell\u2019 (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>). It mitigates technical drop-out in single-cell gene expression profiling by aggregating over multiple cells while simultaneously reducing the influence of cell-type abundance. Hence, the snGEPc representation is suitable for ecotype detection based on cell-state variation, rather than shifts in local cell-type composition alone.<\/p>\n<p>SN similarity network construction<\/p>\n<p>To incorporate the spatial proximity of distinct cell types, a pairwise similarity matrix Ac of dimension m\u2009\u00d7\u2009m is constructed for each matrix Ec (step 2, left; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). In detail, Spatial EcoTyper first performs dimension reduction on matrix Ec to identify the top 20 principal components using the RunPCA function from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e3977\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>. Pairwise similarities between all Ec columns are then calculated as inverted Euclidean distance, yielding matrix Ac for each cell type c. Given the typically large number of SNs, we retained only the top \u03b1 highest similarities in each row and column of Ac, setting all other values to zero to create a sparse matrix. Although \u03b1\u2009=\u200950 was used in this work, we note that our results were robust to a range of empirically tested values of \u03b1 (data not shown). This step maintains key edges in the similarity network and enhances scalability. For any instance in which the given cell type c was not represented in both SNs, the corresponding entry in Ac was assigned as NA.<\/p>\n<p>SN similarity network fusion<\/p>\n<p>Because all SNs are co-registered across cell types, Spatial EcoTyper fuses all Ac matrices into a single similarity matrix A of dimension m\u2009\u00d7\u2009m using SNF (step 2, right; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). Matrix A combines shared patterns of transcriptional covariance across colocalized cell types. To achieve network fusion in practice, we implemented an enhanced version of the SNF function from the SNFtool R package (v.2.3.1), adding support for sparse matrices and missing values while otherwise preserving the original functionality, then we applied this updated function to our set of Ac matrices. We then performed a rank normalization over the columns of the resulting matrix to transform similarity values per column into a standard space. Ranked values per column were subsequently converted to zero minimum and unit maximum.<\/p>\n<p>Spatial clustering and cluster profiling<\/p>\n<p>Given the fused similarity matrix A, Spatial EcoTyper groups SNs into clusters (step 3; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>), which will become candidates for SEs when considered across multiple samples as described in section \u2018Identification of conserved SEs\u2019 below. Clustering each input sample prior to multisample SE discovery serves two related purposes. First, it reduces the dimensionality of the data by grouping SNs with similar spatial covariance patterns. Second, it simplifies cross-sample integration by de-noising the data and minimizing drop-out. To cluster matrix A, we leveraged standard processing procedures optimized within Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4076\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>, sequentially applying RunPCA, FindNeighbors and FindClusters (Louvain) functions. For single-sample analyses, the resulting spatial clusters represent sample-level SEs. For integrative analysis across samples, a higher clustering resolution is recommended to enhance robustness to parameter variation (section\u00a0\u2018Louvain clustering resolution\u2019 below). In this work, we selected a resolution of 30, reducing the dimensionality from tens of thousands of individual SNs to hundreds of SN clusters per sample. Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a> shows the robustness of SE discovery to different clustering resolutions (1\u201350).<\/p>\n<p>In practice, SNs in pre-annotated domains can be balanced to ensure equal representation before clustering. To obtain an equal number of SNs from tumour and adjacent stromal regions in this work, we uniformly downsampled the one with more SNs (for example, tumour) before integrative analysis.<\/p>\n<p>Identification of conserved SEs<\/p>\n<p>Beyond sample-level SE analysis, a key strength of Spatial EcoTyper lies in its ability to identify SEs conserved across a variety of conditions, such as samples, patients and cancer types. To identify such SEs, Spatial EcoTyper uses a variant of the sample-level process described above (section\u00a0\u2018Determination of sample-level spatial clusters\u2019), using SN clusters rather than SNs as the fundamental units (steps 4\u20139; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>).<\/p>\n<p>Assembly of cell-type-specific spatial cluster gene expression profiles<\/p>\n<p>Following clustering of SNs in a given input sample (section \u2018Spatial clustering and cluster profiling\u2019 above), each cell is assigned to the cluster of its spatially nearest SN. To minimize batch effects across samples, row-based standardization is then applied to each single-cell GEP, normalizing gene expression to zero mean and unit variance per gene. Spatial EcoTyper then computes the average cell-type-specific GEP for each cell type c and SN cluster, referred to as ccGEPc.<\/p>\n<p>Once defined, ccGEPc vectors are aggregated per cell type c for a given input sample into matrix E\u2032c, with g genes (rows) and s SN clusters (columns), with genes restricted to those with non-zero expression in at least some SN in each sample (step 4; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>). Importantly, all E\u2032c matrices are co-registered across cell types, with any given SN cluster occupying the same column index across E\u2032c matrices. Moreover, to ensure sufficient representation and well-defined computations per sample, we require a minimum of three SN clusters containing the cell type c for inclusion of a sample into each E\u2032c. In preparation for cross-sample integration, the feature space of each E\u2032c is reduced to the top 200 variable genes (by default), where highly variable genes per matrix are computed according to their rank product of variances across all input samples. In other words, for each cell type c, the variance of each gene across SN cluster ccGEPc vectors is computed per sample, with genes then assigned a rank by variance. Ranks are then aggregated across samples by geometric mean, and the top highly variable genes are selected per E\u2032c from the result.<\/p>\n<p>Cross-sample SN similarity network construction<\/p>\n<p>When E\u2032c matrices have been created for all input samples (step 5; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>), they are concatenated column-wise across samples yielding E*c, stratified by cell type c (step 6; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>). Similarity networks are then computed by Spearman correlation over the columns of E*c for each cell type c, yielding a set of c pairwise similarity matrices A*c describing the similarity of E*c columns across sample-level SN clusters (step 7; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>). For any pairwise comparison of SN clusters in which cell type c is not represented in both clusters, the corresponding entry in A*c is assigned NA. To minimize any remaining batch effects, Spatial EcoTyper standardizes similarity matrices A*c by performing rank normalization independently on each submatrix \\({{\\bf{A}}}_{{ij}}^{* c}\\), which represents the similarities of SN clusters between samples i and j. The normalization is performed by converting the non-NA entries of each column in \\({{\\bf{A}}}_{{ij}}^{* c}\\) to ranks and rescaling the ranks to the unit interval.<\/p>\n<p>Cross-sample SN similarity network fusion<\/p>\n<p>In step 8 (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>), Spatial EcoTyper fuses all A*c matrices across cell types into a single similarity matrix A* using the enhanced SNF function as described above in \u2018SN similarity network fusion\u2019. The resulting multisample matrix encodes the conservation of spatial community structures across cell types and SNs.<\/p>\n<p>Clustering of sample-level SN clusters into SEs<\/p>\n<p>In the final step, to group sample-level SN clusters into cross-sample SEs, NMF clustering is applied to A*, with the number of clusters (rank) set according to the following procedure (step 9; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>). In this work, NMF clustering of A* was tested for ranks ranging from 2 to 50, with 50 runs per rank using the Brunet method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Brunet, J.-P., Tamayo, P., Golub, T. R. &amp; Mesirov, J. P. Metagenes and molecular pattern discovery using matrix factorization. Proc. Natl Acad. Sci. USA 101, 4164&#x2013;4169 (2004).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR82\" id=\"ref-link-section-d74592869e4387\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a>, with optimal threshold selected as the highest rank for which the cophenetic coefficient exceeded 0.95 and subsequently showed the greatest drop (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4b<\/a>). NMF results derived from the selected rank, here identified as 11, were used to group sample-level SN clusters and corresponding SNs and single cells into SEs. This number was further reduced to nine final SEs by excluding candidate ecotypes that were devoid of SE-specific cell states (section \u2018Discovery of conserved SE-specific cell states\u2019\u00a0below) and that did not exhibit maximal within-cluster similarity. Similarity between two clusters was calculated as the average value of the block of A* corresponding to rows of the first cluster and columns of the second. Notably, NMF has well-established performance characteristics for robustly clustering high-dimensional genomic data encompassing hundreds to thousands of data points<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e4398\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 70\" title=\"Barkley, D. et al. Cancer cell states recur across tumor types and form specific interactions with the tumor microenvironment. Nat. Genet. 54, 1192&#x2013;1201 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR70\" id=\"ref-link-section-d74592869e4401\" rel=\"nofollow noopener\" target=\"_blank\">70<\/a>. However, for step 3 of the multi-sample workflow (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>), we used the more efficient Louvain clustering (Seurat), owing to the large SNF matrices arising from single-sample analysis. For robustness testing, see Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>.<\/p>\n<p>Assembly of cell-type-specific SE gene expression profiles<\/p>\n<p>Having assigned single cells to SEs, Spatial EcoTyper then determines SE cell-state gene expression profiles (csGEPs). For each cell type c, NMF is performed on single-cell GEP Gc and corresponding SE label matrix Hc to derive csGEPs. Here, Gc represents a gene-by-cell matrix for each cell type c. To construct Gc, GEPs are normalized with NormalizeData from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4450\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>, standardized to zero mean and unit variance per gene in each sample, and then posneg transformed as described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e4454\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>. To ensure balanced representation, an equal number of cells (at least 300, and up to 5,000 cells) are randomly selected from each sample-SE pair. Owing to the computational constraints of NMF, a maximum of 25,000 cells is selected through random down-sampling. GEPs of the selected cells are concatenated column-wise into csGEP matrix Gc. The SE label matrix Hc is a binary cell by SE matrix indicating SE membership for each cell in Gc. NMF is then applied to solve the equation:<\/p>\n<p>$${{G}^{c}=W}^{c}\\times {H}^{c},$$<\/p>\n<p>where Wc represents the basis matrix containing csGEPs. To refine csGEPs for cell-state recovery, the top 50 genes are chosen per SE based on the largest positive delta compared with the second-highest expression across SEs. Each basis matrix is then reduced to the selected genes. These refined basis matrices enable the recovery of SEs and their cell states from independent data using NMF (section \u2018Recovery of SE-specific cell states\u2019 below).<\/p>\n<p>Discovery of conserved SE-specific cell states<\/p>\n<p>Although SEs are derived from spatial covariation in cell states across cell types, shared across samples, not every cell state associated with an SE need be specific to that SE. To identify and validate cell states specifically enriched in each SE and conserved across discovery samples, we performed leave-one-sample-out cross-validation (LOOCV), repeating the csGEP construction as described above (\u2018Assembly of cell-type-specific SE gene expression profiles\u2019\u00a0above) for each training fold. We then used the resulting NMF basis matrices to predict cell-state labels on the held-out sample. For label assignment, NMF prediction output matrices Hc were standardized to unit sum per column, yielding a probability matrix encoding the probability of each single cell being localized in each SE. Cells were then assigned to the SE-associated cell state with the highest probability. For each LOOCV iteration, the enrichment of each cell state in each SE was assessed by its ability to correctly assign cells to that SE using an F1 score.<\/p>\n<p>Because the csGEP construction involves subsampling of cells, this LOOCV process was repeated 20 times to ensure robustness, with F1 scores averaged across all repetitions. Cell states were considered specific to an SE only if the associated F1 score exceeded the second highest F1 score for the cell state across other SEs by at least 0.1. Otherwise, the cell state was deemed either broadly distributed across multiple SEs or not conserved across samples. Using this approach, 38 SE-specific cell states were identified, specific to nine SEs (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2f<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). Two SEs identified initially were excluded from further analysis owing to a lack of specific cell states, and the remaining SEs were renumbered accordingly from 1 to 9 based on their average distance across discovery samples to the tumour margin (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2d<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4c<\/a>).<\/p>\n<p>Recovery of SE-specific cell states<\/p>\n<p>After identifying conserved SE-specific cell states through the above LOOCV process, we used all discovery-cohort samples to prepare an ensemble basis matrix W*c for each cell type c. Specifically, for each cell type, we repeated the process described\u00a0above in \u2018Assembly of cell-type-specific SE gene expression profiles\u2019 50 times and then averaged the resulting basis matrices to produce W*c. For feature selection, genes showing the highest and most specific expression in each cell state were identified from each basis matrix as described, and then genes that were selected for the same cell state in more than half of the repetitions were retained in the ensemble matrix.<\/p>\n<p>The resulting ensemble matrices are a core component of the Spatial EcoTyper framework and can be used to recover SE-specific cell states from external single-cell-scale transcriptomics datasets using NMF, as described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e4601\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>. To recover SE-specific cell states from a query dataset, NMF is applied with single-cell-scale gene expression matrix Gc and the ensemble matrix W*c as input, to yield a probability matrix Hc denoting the probability of each cell belonging to each SE. Cells are then assigned to SE-specific cell states when the prediction probability exceeds 0.6, otherwise they are designated to a null class, referred to as non-SE. In practice, single-cell-scale GEPs in Gc should be normalized (to counts per million (CPM), TPM or by SCTransform<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4631\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>, as appropriate) and then scaled to zero mean and unit variance per gene.<\/p>\n<p>Deconvolution of SEs from bulk RNA-seq<\/p>\n<p>To enable profiling of SEs in bulk expression data, the Spatial EcoTyper framework includes an NMF model trained over simulated bulk RNA-seq prepared by aggregation of scRNA-seq data into pseudo-bulk mixtures for which ground-truth SE proportions are known.<\/p>\n<p>Construction of pseudo-bulk mixtures<\/p>\n<p>Previously described publicly available scRNA-seq data (section\u00a0\u2018scRNA-seq\u2019 above) from ten cancer types were used to create pseudo-bulk mixtures. First, cells were annotated as described in section\u00a0\u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019 below, labelled according to the parent SE of their assigned cell state or, if unassigned, designated non-SE, for a total of ten label classes. The fractional composition of each pseudo-bulk was generated by random sampling of a value per label class from the Gaussian distribution N(\u03bc\u2009=\u20092, \u03c3\u2009=\u20091), with negative values set to zero and the resulting values normalized to unit sum.<\/p>\n<p>Pseudo-bulks were assembled separately by cancer type, with GEPs constructed by aggregating 1,000 cells randomly selected to satisfy the predefined fractions of cell states. For UMI- and plate-seq-based data, raw counts and TPM values were respectively summed across selected cells. We generated 100 pseudo-bulks per cancer type, and the resulting GEPs were normalized using the NormalizeData function from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4662\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>. To mitigate cancer type and batch differences, GEPs were further normalized to zero mean and unit variance per gene in each cancer type.<\/p>\n<p>NMF model training for bulk deconvolution<\/p>\n<p>The resulting profiles were concatenated into gene by sample matrix EP, with genes limited to those detected across all cancer types, and pseudo-bulk SE fractional abundances were encoded in sample by label matrix HP. From these, a basis matrix was derived by application of NMF followed by feature selection as described above (section\u00a0\u2018Assembly of cell-type-specific SE gene expression profiles\u2019). The resulting basis matrix WB constitutes another core component of the Spatial EcoTyper framework and can be used to deconvolve SE fractional abundances from bulk gene expression data. For a given bulk expression dataset, predictions are performed by NMF as described in section\u00a0\u2018Discovery of conserved SE-specific cell states\u2019, excluding the final classification step to yield HP, in which the values represent SE abundances across input bulk samples. In practice, to perform deconvolution, input data should be normalized to TPM or CPM as appropriate, log2-adjusted and then normalized across samples to zero mean and unit variance per gene.<\/p>\n<p>Spatial EcoTyper discovery cohort<\/p>\n<p>MERSCOPE samples were selected for SE discovery owing to their high spatial resolution and the availability of uniformly processed samples across multiple cancer types (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). Before analysis, MERSCOPE samples were preprocessed as described in section\u00a0\u2018MERSCOPE\u2019 above, then standardized using NormalizeData from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4713\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>. For SE discovery and characterization, we focused on nine main TME cell types with strong representation across cancer types and tumour samples: B cells, plasma cells, CD4+ T cells, CD8+ T cells, NK cells, macrophages, dendritic cells, fibroblasts and endothelial cells. Malignant cells were not included owing to significant differences across tumour types. Other TME cell types, such as smooth muscle cells, pericytes and neutrophils, were not confidently detected in the MERSCOPE dataset, probably because of limitations in transcript capture inherent to MERSCOPE and the 500-gene panel that we analysed.<\/p>\n<p>To capture spatial microenvironments from both tumour and adjacent stroma (annotated as described in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>), we selected samples in which each region included more than 5% of the total TME (immune and\/or stromal) cells, yielding two melanomas, two colon cancer and two liver cancer samples, and one breast cancer, one prostate cancer and one ovarian cancer sample (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">3f<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). Five of these samples, each from a different cancer type, were used for SE discovery (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">3f<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>).<\/p>\n<p>Selection of Spatial EcoTyper parametersSN radius<\/p>\n<p>When applying Spatial EcoTyper to individual samples, we consistently observed a spatial gradient resembling the physical distance of SNs to the tumour margin (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a>). To assess robustness, Spatial EcoTyper analyses were performed with ten different SN radii, ranging from 10\u2009\u03bcm to 100\u2009\u03bcm, on a MERSCOPE melanoma sample (melanoma 1). For each SN radius, the relationship between gene expression similarity and physical distance of SNs to the tumour margin was evaluated. To do this, we used the procedure described in \u2018Cells, meta-cells, and Spatial EcoTyper embeddings vs. distance to the margin\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>, with the exception that PCA was applied to the spatial embedding produced by Spatial EcoTyper (step 2; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>) and SNs (rather than single cells) were used as the unit of analysis. To strike a balance between SN granularity and correlation with distance to the margin, a radius of 50\u2009\u00b5m was selected for SE discovery (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">3g<\/a>).<\/p>\n<p>Louvain clustering resolution<\/p>\n<p>A key parameter in the Spatial EcoTyper discovery pipeline (step 3; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>) is the resolution for Louvain clustering, which groups SNs into clusters in each sample. To ensure robustness, 11 different resolutions ranging from 1 to 50 were tested, and all resulting spatial clusters were grouped into ten clusters following the multisample discovery pipeline. The similarity between clusters derived at different resolutions was evaluated using the average adjusted Rand index (ARI), comparing each resolution with the others (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>). The discovered clusters showed high overall consistency, with results being more stable when a resolution higher than 15 was used. The resolution of 30, which had the highest average ARI, was selected for SE discovery in our analysis (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>).<\/p>\n<p>Recovery of cell states and SEs in ST and scRNA-seq validation datasetsSingle-cell-scale ST recovery<\/p>\n<p>To validate SEs and their associated cell states, we analysed nine samples profiled by MERSCOPE (five discovery samples and four held-out samples) and 12 held-out samples profiled by different ST platforms: Xenium V1 (n\u2009=\u20095), Xenium Prime (n\u2009=\u20094) and Visium HD (n\u2009=\u20093) (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). The latter, drawn from publicly available data described in section\u00a0\u2018Data collection and processing\u2019 above, were selected for consistency with the TME content threshold required for the MERSCOPE discovery cohort (section\u00a0\u2018Spatial EcoTyper discovery cohort\u2019 above). In all held-out samples, SE-specific cell states were recovered using the approach described above\u00a0in section\u00a0\u2018Recovery of SE-specific cell states\u2019, assigning single cells (or 8-\u00b5m2 bins from Visium HD) to either non-SE or SE-specific cell states, which allowed for further grouping into the respective SEs. The same procedure was also applied to the MERSCOPE discovery cohort but with LOOCV to avoid overfitting (section \u2018Discovery of conserved SE-specific cell states\u2019 above).<\/p>\n<p>Bulk ST recovery<\/p>\n<p>For the analysis presented in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6j<\/a>, 26 Visium and 48 legacy ST samples were selected, each containing at least five spots located more than 500\u2009\u00b5m away from the tumour margin (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>). Spatial spots were normalized to zero mean and unit variance per gene across all spots in each sample. We applied the SE cell-state recovery models (section\u00a0\u2018Recovery of SE-specific cell states\u2019\u00a0above) to obtain an H*c matrix for each cell type c, and then averaged the matrices across cell types to estimate relative SE levels across spots. Each spot was then assigned to the dominant SE.<\/p>\n<p>Single-cell RNA-seq recovery<\/p>\n<p>For the analyses presented in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7b<\/a>, we queried SE content in scRNA-seq profiles from 144 tumour samples spanning ten cancer types (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). For each cell type from each carcinoma type, the scRNA-seq data were normalized using SCTransform from Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4846\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>, and log2 TPM data were used for melanoma Smart-seq2 profiles. Cell-state recovery was then performed as described above (\u2018Recovery of SE-specific cell states\u2019), and the abundance of cell state i of parent cell type c (for example, CD8+ T cells) in sample s was then determined as the fraction of cells assigned to state i out of the total cells of cell type c in sample s. We repeated the above process for scRNA-seq profiles of 64 brain metastases from melanoma and five types of carcinoma<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Xing, X. et al. Pan-cancer human brain metastases atlas at single-cell resolution. Cancer Cell 43, 1242&#x2013;1260 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR36\" id=\"ref-link-section-d74592869e4874\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>, using NormalizeData with Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e4878\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a> before SE recovery (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>).<\/p>\n<p>Metrics and analyses for SE recovery in ST and scRNA-seq dataCell-state colocalization in single-cell-scale ST data<\/p>\n<p>The spatial colocalization patterns of SE-specific cell states were assessed using single-cell-scale ST data from four platforms, with cell states recovered from each sample as described above (\u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). For each single cell (or 8-\u00b5m bin for Visium HD), the fractional abundances of neighbouring cell states within a radius of 50\u2009\u00b5m were determined, resulting in an N\u2009\u00d7\u2009S matrix, F, where Fij denotes the fraction of cells within a 50-\u00b5m radius of cell Ni that are of state Sj. Single-cell-level fractions were subsequently averaged in each cell state, producing an S\u2009\u00d7\u2009S colocalization matrix, L, where Lij represents the average fractional abundance of cell state Si near cell state Sj. To control for biases, cell-state assignments were shuffled 10,000 times and colocalization matrices were recomputed, yielding 10,000 random colocalization matrices, Lrand. The colocalization matrix L was then normalized by subtracting the average of the Lrand matrices and dividing by their standard deviation for each element:<\/p>\n<p>$${L}_{ij}^{{\\prime} }=\\frac{{L}_{ij}-\\,{\\mu }_{{L}_{ij}^{{\\rm{r}}{\\rm{a}}{\\rm{n}}{\\rm{d}}}}}{{{\\rm{\\sigma }}}_{{L}_{ij}^{{\\rm{r}}{\\rm{a}}{\\rm{n}}{\\rm{d}}}}}.$$<\/p>\n<p>This resulted in a matrix L\u2032, where Lij\u2032 represents the colocalization index between cell state Si and cell state Sj. Finally, the colocalization indexes from multiple samples in each dataset were integrated using Stouffer\u2019s approach<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Stouffer, S. A., Suchman, E. A., Devinney, L. C., Star, S. A. &amp; Williams, R. M. Jr. The American Soldier: Adjustment during Army Life. (Studies in Social Psychology in World War II) Vol. 1 (Princeton Univ. Press, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR80\" id=\"ref-link-section-d74592869e5173\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>, with L\u2032 capped at an absolute value of 5 per sample to prevent any single sample from disproportionately influencing the meta-analysis (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6c\u2013g<\/a>).<\/p>\n<p>Cell-state co-association in scRNA-seq data<\/p>\n<p>Cell-state abundances were determined in scRNA-seq tumour atlases as described above\u00a0in \u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019. Abundances were computed under four schemes: (i) including all SE and non-SE states; (ii) excluding non-SE states; third, the same as (i) except treating zero abundance as missing values (NA); and (iv) the same as (ii) except treating zero abundance as NA. Pairwise Pearson correlations between cell states were then calculated across all scRNA-seq samples for each abundance matrix, using the cor function in R with pairwise complete observations. The final co-association values were obtained by averaging the correlations across the four schemes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7a\u2013c<\/a>).<\/p>\n<p>Significance of cell-state colocalization and co-association<\/p>\n<p>Permutation experiments were done to assess the significance of cell-state cooccurrence indices, whether for colocalization in ST data (L\u2032 from \u2018Cell-state colocalization in single-cell-scale ST data\u2019 above) or for co-associations in scRNA-seq data (\u2018Cell-state co-association in scRNA-seq data\u2019 above). Let square matrix C represent all pairwise co-occurrence indices between SE cell states. Let \u0398w represent the average of all co-occurrence indices in C for cell states within SE class w. \u0398w was compared with 10,000 corresponding scores \\({\\theta }_{w}^{{\\rm{rand}}}\\) obtained by randomly shuffling the order of all columns in C, then determining the average co-occurrence index for all cell-state indices corresponding to SE w. The mean co-occurrence score \u0398w was then normalized by subtracting the average of \\({\\theta }_{w}^{{\\rm{rand}}}\\) and dividing by its standard deviation, yielding a two-sided z-score. For scRNA-seq co-association analyses, the z-score was directly converted into a P-value, and the process was repeated for each SE (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7b,c<\/a>). For ST colocalization analyses, each ST sample was analysed individually. Stouffer\u2019s method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Stouffer, S. A., Suchman, E. A., Devinney, L. C., Star, S. A. &amp; Williams, R. M. Jr. The American Soldier: Adjustment during Army Life. (Studies in Social Psychology in World War II) Vol. 1 (Princeton Univ. Press, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR80\" id=\"ref-link-section-d74592869e5316\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a> was then used to aggregate SE-specific z-scores across all ST samples in a dataset, resulting in a meta-z-score for each SE, which was directly converted into a P-value (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6d\u2013f<\/a>). To incorporate SE2, which is composed of a single cell state, colocalization testing was applied to individual SE2 cells. Otherwise, self-comparisons of cell states were excluded from significance testing.<\/p>\n<p>Spatial autocorrelation<\/p>\n<p>The spatial coherence of SE-specific cell states was evaluated using Moran\u2019s I across 21 single-cell-scale ST datasets, with cell states recovered independently for each sample using Spatial EcoTyper (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6j<\/a>). We constructed a k-nearest-neighbour graph (k\u2009=\u20093) using spdep (v.1.3.11)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 83\" title=\"Pebesma, E. &amp; Bivand, R. Spatial Data Science: With Applications in R (Chapman and Hall\/CRC, 2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR83\" id=\"ref-link-section-d74592869e5353\" rel=\"nofollow noopener\" target=\"_blank\">83<\/a> over all annotated TME cells. We converted the resulting graph into a spatial weights matrix using the nb2listw function (with default parameters) and then calculated Moran\u2019s I for each SE class using the moran function, where cells belonging to the SE were encoded as 1 and all others as 0.<\/p>\n<p>To control for bias, we performed 1,000 permutation experiments in which SE labels were randomly shuffled within each cell type. Moran\u2019s I was recalculated for each permutation to generate null distributions for every SE. Observed Moran\u2019s I values were then normalized into z-scores by subtracting the mean of the null distribution and dividing by its standard deviation.<\/p>\n<p>Distance to tumour margin<\/p>\n<p>Following SE recovery from ST data as described above\u00a0in section\u00a0\u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019, the distance of each SE to the tumour margin (in micrometres) was computed by first averaging the Euclidean distance to the nearest tumour margin of all SNs (single-cell-scale ST) or spots (bulk ST) assigned to each SE in each sample, then averaging the resulting quantities by SE across samples in each cohort (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6i,j<\/a>). Positive and negative distances were used for cells and spots localized to the tumour region and adjacent stroma, respectively (see, for example, Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1c<\/a>). Finally, for each cohort, the consistency between expected distances (SE-specific distances to the tumour margin derived from the MERSCOPE discovery cohort) and predicted distances was evaluated using Pearson correlation (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">6i,j<\/a>).<\/p>\n<p>Characterization of SEs and associated cell statesIdentification of SE\u00a0cell-state markers<\/p>\n<p>To identify SE-specific cell-state markers, we analysed scRNA-seq data from 144 tumours and ten cancer types (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>), with all cells grouped into SE-specific cell states or the non-SE null class. The scRNA-seq data were normalized as described in section \u2018Differential expression analysis\u2019 above. Differential expression analysis was performed by comparing each cell state with all of the other cells of the same cell type using the wilcoxauc function from the presto package (v.1.0.0; <a href=\"https:\/\/github.com\/immunogenomics\/presto\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/immunogenomics\/presto<\/a>)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"Korsunsky, I., Nathan, A., Millard, N. &amp; Raychaudhuri, S. Presto scales Wilcoxon and auROC analyses to millions of observations. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/653253&#010;                &#010;               (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR79\" id=\"ref-link-section-d74592869e5414\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>. LFCs were extracted from each cancer type and then aggregated across the ten cancer types by median, yielding pan-cancer LFCs.<\/p>\n<p>To identify the markers most specific to each cell state i, we selected genes whose pan-cancer LFC in i was at least 0.1 higher than in any other state. Among these, the top 30 genes by pan-cancer LFC were considered to be marker genes (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2g<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). Markers for all SE cell types annotated in at least five cancer types in the combined scRNA-seq tumour atlas were determined (all except B and NK cells).<\/p>\n<p>Reproducibility of cell-state markers<\/p>\n<p>To assess the reproducibility of cell-state marker genes, we parcelled all ten scRNA-seq atlases into discovery (n\u2009=\u20095 cancer types corresponding to those used for the MERSCOPE discovery cohort) and validation (n\u2009=\u20095 remaining cancer types) cohorts, each with non-overlapping cancer types. Marker genes identified in the discovery cohort were evaluated in the validation cohort (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7d<\/a>) and compared with markers derived from the validation cohort or all ten cancer types using the Jaccard similarity index (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7e<\/a>).<\/p>\n<p>Annotation of cell states<\/p>\n<p>SE-specific cell states (n\u2009=\u200938) were annotated based on top markers and 135 previously reported reference cell states (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). Specifically, we assessed similarity to each of the 135 reference states by computing enrichment scores using AddModuleScore in Seurat (v.4.3.0)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573&#x2013;3587 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR67\" id=\"ref-link-section-d74592869e5467\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>. For each reference state, we then averaged enrichment scores across all cells of the state and tested whether the mean score was significantly higher than that of randomly sampled cells by permutation testing over 1,000 iterations. Of the 38 SE-specific states, 18 showed significant overlap with at least one reference state and were annotated with the most associated reference state by significance (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). To augment these assignments, all 38 SE cell states were also annotated based on the corresponding marker gene with the highest LFC compared with other cell states of the same cell type (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2g<\/a>, Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>).<\/p>\n<p>Identification of SE consensus markers<\/p>\n<p>To identify SE-specific genes with conservation across cell types, termed consensus markers (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2h<\/a>), the top 1,000 markers with positive pan-cancer LFCs\u2014or all positive markers if fewer than 1,000 were available\u2014were selected for each SE cell state from the analysis of scRNA-seq data described above (\u2018Identification of SE cell-state markers\u2019). Consensus SE markers were then defined as genes with at least 80% conservation across all evaluable cell states in a SE (equivalent to 100% conservation for ecotypes with fewer than five states) for a minimum of 20 markers per SE. For SE2, which comprises a single-cell state, we limited consensus markers to those with statistically significant conservation (Q\u2009&lt;\u20090.05) in at least three cancer types. To eliminate overlap among consensus markers, genes associated with multiple SEs were assigned to the SE with the highest number of significantly conserved cancer types. In cases where markers overlapped between SE2 and other SEs, genes were preferentially assigned to non-SE2 states. Given that SE3 had fewer than 20 genes following these steps (n\u2009=\u200916), we augmented it by including genes at a relaxed cell-type-conservation threshold of 60%, selected in order of decreasing conservation across cancer types, until the 20-marker minimum was satisfied. Consensus markers and normalized expression values are provided in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a> (see also\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>).<\/p>\n<p>Biological pathways associated with SE consensus markers<\/p>\n<p>To identify biological pathways associated with SE consensus markers, we performed overlap analysis using the enricher function from the clusterProfiler (v.4.14.6) R package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 84\" title=\"Yu, G. Thirteen years of clusterProfiler. Innovation 5, 100722 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR84\" id=\"ref-link-section-d74592869e5513\" rel=\"nofollow noopener\" target=\"_blank\">84<\/a>. Consensus marker sets were individually evaluated against hallmark (H) and biological process (C5:BP) gene sets from MSigDB<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 85\" title=\"Liberzon, A. et al. The Molecular Signatures Database (MSigDB) hallmark gene set collection. Cell Syst. 1, 417&#x2013;425 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR85\" id=\"ref-link-section-d74592869e5517\" rel=\"nofollow noopener\" target=\"_blank\">85<\/a>. Pathways with significant overlap (Q\u2009&lt;\u20090.1) were retained, and for each SE, pathways showing the strongest overlap relative to other SEs were selected (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2h<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7h<\/a>).<\/p>\n<p>Association between SEs and carcinoma ecotypes<\/p>\n<p>To study the relationship between SEs and previously defined CEs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e5538\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>, SEs and CEs were recovered from the same scRNA-seq data across ten cancer types (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>) using SE-specific and previously published<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e5545\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a> recovery methods, respectively. The fraction of cells in each SE i that were also assigned to CE j was computed for each dataset, resulting in an overlap matrix O with rows representing nine SEs and columns representing nine CEs (excluding CE7 because of its low validation rate in a previous study<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e5559\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>). To control for potential biases from different abundances, permutation experiments were performed by shuffling the cell-state assignments 10,000 times. In each iteration, the matrix O was recomputed, yielding \\({O}_{i,j}^{{\\rm{rand}}}\\). The matrix O was then normalized using the mean and variance of \\({O}_{i,j}^{{\\rm{rand}}}\\) (as described in section\u00a0\u2018Cell-state colocalization\u2019\u00a0above), producing a normalized matrix O\u2032, where \\({O}_{i,j}^{{\\prime} }\\) represents the overlap index between SE i and CE j. Finally, the overlap indexes across the ten cancer types were aggregated using Stouffer\u2019s method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Stouffer, S. A., Suchman, E. A., Devinney, L. C., Star, S. A. &amp; Williams, R. M. Jr. The American Soldier: Adjustment during Army Life. (Studies in Social Psychology in World War II) Vol. 1 (Princeton Univ. Press, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR80\" id=\"ref-link-section-d74592869e5686\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a> (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">7j<\/a>).<\/p>\n<p>Validation of SE deconvolutionCross-validation over training pseudo-bulks<\/p>\n<p>To evaluate the Spatial EcoTyper deconvolution model (section\u00a0\u2018Deconvolution of SEs from bulk RNA-seq\u2019 above), we first applied a LOOCV procedure in which we trained the NMF model on pseudo-bulk GEPs from nine cancer types (n\u2009=\u2009900 mixtures) and applied the trained model to pseudo-bulk GEPs from the remaining cancer type (n\u2009=\u2009100 mixtures; see\u00a0section \u2018Construction of pseudo-bulk mixtures\u2019\u00a0above). Consistency between predicted and ground-truth SE abundances was assessed by Pearson correlation (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3a<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8a\u2013c<\/a> and \u2018Benchmarking of SE deconvolution\u2018 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>).<\/p>\n<p>Paired bulk RNA-seq and scRNA-seq<\/p>\n<p>We further evaluated the Spatial EcoTyper deconvolution model using paired scRNA-seq and bulk RNA-seq data from cohort 1 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3a,d<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8d<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>). SE cell states were first recovered from scRNA-seq data (section \u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019 above) and SE abundances defined as the number of cells assigned to each SE over total evaluable cells per sample. Because mitochondrial quality-control filtering (section\u00a0\u2018Data collection and processing\u2019 above) disproportionately removed cancer cells from a minority of samples, cancer and TME abundances were rescaled to match their proportions in the mitochondrial-unfiltered data, and SE proportions were adjusted accordingly. Next, SE abundances were inferred from bulk RNA-seq using the Spatial EcoTyper deconvolution model (section \u2018NMF model training for bulk deconvolution\u2019 above). Owing to the limited number of samples, which could bias the centring and unit variance normalization of gene expression required for SE deconvolution, we combined cohort 1 and TCGA RNA-seq data from matched cancer types, removed batch effects between the two datasets using the Combat function from the sva R package (v.3.46.0) with default parameters<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 86\" title=\"Johnson, W. E., Li, C. &amp; Rabinovic, A. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics 8, 118&#x2013;127 (2007).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR86\" id=\"ref-link-section-d74592869e5738\" rel=\"nofollow noopener\" target=\"_blank\">86<\/a>, and then normalized the batch-corrected gene expression to zero mean and unit variance per gene across all samples. This process was conducted for melanoma and colon cancer separately, and for intact and digested bulk RNA-seq datasets separately (\u2018Bulk RNA sequencing\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>). The resulting data were used to infer SE abundances with Spatial EcoTyper (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8e<\/a> and \u2018Benchmarking of SE deconvolution\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>).<\/p>\n<p>Paired bulk RNA-seq and Visium ST<\/p>\n<p>We downloaded bulk RNA-seq data as FASTQ files (n\u2009=\u200954 samples) and preprocessed Visium ST profiles (n\u2009=\u2009103 samples) from 47 patients spanning five carcinoma types from the HTAN<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Rozenblatt-Rosen, O. et al. The Human Tumor Atlas Network: charting tumor transitions across space and time at single-cell resolution. Cell 181, 236&#x2013;249 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR44\" id=\"ref-link-section-d74592869e5766\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a>. These data were processed and subjected to quality control as described in \u2018Bulk RNA sequencing\u2019 and \u2018Paired bulk RNA-seq and ST quality control\u2019, respectively, in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>, resulting in matched pairs from 42 of 47 patients, comprising 46 bulk RNA-seq and 88 Visium samples. Next, bulk RNA-seq data (log2 TPM) were scaled to unit variance per gene across samples for input to the Spatial EcoTyper deconvolution model (section \u2018NMF model training for bulk deconvolution\u2019\u00a0above). The log2 CPM data from each Visium sample were scaled to mean of zero and unit variance per gene across all spots. Deconvolution was then performed for each sample separately to determine SE abundances across Visium spots. To obtain sample-level SE abundances accounting for geographic variation in cell density, we estimated cell counts per spot using CytoSPACE (v.1.0.3)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Vahid, M. R. et al. High-resolution alignment of single-cell and spatial transcriptomes with CytoSPACE. Nat. Biotechnol. 41, 1543&#x2013;1548 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR25\" id=\"ref-link-section-d74592869e5778\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>. We then used these count estimates to compute a weighted average of SE levels across all spots and renormalized the resulting SE levels to unit sum in each sample. We averaged SE levels across samples per modality for each patient. The resulting SE abundances from Visium and matched bulk RNA-seq were then compared (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3b\u2013d<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8h,i<\/a>). We also repeated this analysis, replacing bulk RNA-seq data with pseudo-bulk profiles constructed from Visium samples as described in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>, but without batch correction (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig13\" rel=\"nofollow noopener\" target=\"_blank\">8h<\/a>).<\/p>\n<p>Paired Visium and single-cell-scale ST<\/p>\n<p>To assess whether spot-level deconvolution from bulk Visium data is consistent with single-cell-scale SE recovery (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3e\u2013f<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig14\" rel=\"nofollow noopener\" target=\"_blank\">9a,b<\/a>), we prospectively generated paired Visium data (quality control and processing as described in \u201810x Visium (standard)\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>) and MERSCOPE data (\u2018Vizgen MERSCOPE\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>) from adjacent melanoma sections (melanoma 3 from patient WU2109; Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>). We also downloaded matched Visium and Visium HD (8-\u03bcm2 bins) data for adjacent colon cancer sections (\u2018Visium HD, Sample P2 CRC\u2019 and \u2018Visium CytAssist v2, Sample P2 CRC\u2019) from 10x Genomics (<a href=\"https:\/\/www.10xgenomics.com\/platforms\/visium\/product-family\/dataset-human-crc\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.10xgenomics.com\/platforms\/visium\/product-family\/dataset-human-crc<\/a>) and preprocessed them as described in section\u00a0\u2018Data collection and processing\u2019\u00a0above. To co-register adjacent tissue sections profiled by Visium (standard) ST and paired single-cell-scale ST data (MERSCOPE or Visium HD), we manually selected four reference points at the edges of distinct morphological structures visible in both datasets. These references were used to learn a linear affine transformation function, which was subsequently applied to transform all coordinates from Visium into the coordinate space of the paired single-cell-scale ST dataset.<\/p>\n<p>SE abundances were inferred for all Visium spots as described above (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec67\" rel=\"nofollow noopener\" target=\"_blank\">Paired bulk RNA-seq and Visium ST\u2019<\/a>). SEs were recovered from all single-cell-scale ST data as described in section\u00a0\u2018Recovery of cell states and SEs in ST and scRNA-seq validation datasets\u2019\u00a0above. Two strategies were used to overcome potential imprecision in the co-registration procedure. First, SE abundances in single-cell-scale ST data corresponding to each co-registered Visium spot were estimated as the fraction of cells or bins assigned to each SE within a SN of 50\u2009\u00b5m radius, requiring at least five cells or bins for robustness. Second, SE abundances in both datasets were smoothed by averaging across each co-registered spot and its six nearest neighbours. Non-SE cells, comprising cancer cells, non-SE TME cells and low-confidence cells (for example, because of limited gene detection; section \u2018Recovery of SE-specific cell states\u2019\u00a0above), were excluded from analysis.<\/p>\n<p>Concordance between platforms was determined by Spearman correlation, adjusting for background dependencies between SE levels in the paired single-cell-scale ST sample (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3e\u2013f<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig14\" rel=\"nofollow noopener\" target=\"_blank\">9a,b<\/a>). To do this, pairwise Spearman correlations were computed between the levels of each Visium-derived SE i and each single-cell-scale-ST-derived SE j, conditioning on SE i in the single-cell-scale ST data, for all non-matching SE pairs, using the pcor function in the R package ppcor (v.1.1)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 87\" title=\"Kim, S. ppcor: An R package for a fast calculation to semi-partial correlation coefficients. Commun. Stat. Appl. Methods 22, 665&#x2013;674 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR87\" id=\"ref-link-section-d74592869e5858\" rel=\"nofollow noopener\" target=\"_blank\">87<\/a>. Direct Spearman correlations were calculated for all matching SE pairs. The P-values of the resulting correlation coefficients were transformed into signed \u2013log10 Q-values indicating the polarity of the correlation following Benjamini\u2013Hochberg correction<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 81\" title=\"Benjamini, Y. &amp; Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. B Methodol. 57, 289&#x2013;300 (1995).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR81\" id=\"ref-link-section-d74592869e5870\" rel=\"nofollow noopener\" target=\"_blank\">81<\/a>.<\/p>\n<p>Large-scale assessment of SE levels in human tumoursOverall survival and pathway analysis<\/p>\n<p>We applied the Spatial EcoTyper deconvolution model (section\u00a0\u2018NMF model training for bulk deconvolution\u2019\u00a0above) to infer SE levels from 7,076 bulk tumour RNA-seq profiles across 17 cancer types from TCGA, including melanoma and 16 carcinomas (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>). SE deconvolution was performed separately for each cancer type using TPM data obtained from the PanCanAtlas (<a href=\"https:\/\/gdc.cancer.gov\/about-data\/publications\/pancanatlas\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/gdc.cancer.gov\/about-data\/publications\/pancanatlas<\/a>), which were log2-adjusted. To investigate SE prognostic associations, a Cox regression analysis was conducted to examine the association between SE abundance and patient overall survival, adjusting for age and sex, using the survival R package (v.3.6.4). This analysis was done separately for each cancer type (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3g<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>). To determine the pan-cancer survival associations of SEs, a meta-analysis was done by combining the resulting z-scores from each SE across all 17 cancer types using Stouffer\u2019s method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Stouffer, S. A., Suchman, E. A., Devinney, L. C., Star, S. A. &amp; Williams, R. M. Jr. The American Soldier: Adjustment during Army Life. (Studies in Social Psychology in World War II) Vol. 1 (Princeton Univ. Press, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR80\" id=\"ref-link-section-d74592869e5909\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>. For clarity, all z-scores and meta z-scores were converted to directional \u2013log10 P-values (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3g<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>). For pathway analysis details (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig14\" rel=\"nofollow noopener\" target=\"_blank\">9d<\/a>), see \u2018Pathways associated with inferred SE abundance in TCGA\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>. For the analysis in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig14\" rel=\"nofollow noopener\" target=\"_blank\">9e<\/a>, MHC levels were computed by first averaging the log2 expression of MHC-I (HLA-A\/B\/C) and MHC-II genes (HLA-D*) separately, then averaging the resulting quantities together. Stromal levels were assessed using ESTIMATE (v.1.0.13)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 88\" title=\"Yoshihara, K. et al. Inferring tumour purity and stromal and immune cell admixture from expression data. Nat. Commun. 4, 2612 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR88\" id=\"ref-link-section-d74592869e5955\" rel=\"nofollow noopener\" target=\"_blank\">88<\/a>.<\/p>\n<p>Associations with immunotherapy response<\/p>\n<p>We obtained publicly available bulk tumour RNA-seq data from patients with\u00a0melanoma and carcinoma treated with ICIs, including anti-PD-1, anti-PD-L1 and combinations of anti-PD-1 and anti-CTLA-4 therapies, after tumour sample collection. All patients were grouped into responders (partial or complete response) and non-responders (stable or progressive disease) based on collected clinical information. To mitigate within-dataset heterogeneity, patients who received prior immunotherapy or chemotherapy were separated into independent datasets. A minimum of five responders and five non-responders was required for each dataset, resulting in 1,249 total patients from 15 datasets from 12 studies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Kim, S. T. et al. Comprehensive molecular characterization of clinical responses to PD-1 inhibition in metastatic gastric cancer. Nat. Med. 24, 1449&#x2013;1458 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR51\" id=\"ref-link-section-d74592869e5967\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Jung, H. et al. DNA methylation loss promotes immune evasion of tumours with high mutation and copy number load. Nat. Commun. 10, 4278 (2019).\" href=\"#ref-CR89\" id=\"ref-link-section-d74592869e5970\">89<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Gide, T. N. et al. Distinct immune cell populations define response to anti-PD-1 monotherapy and anti-PD-1\/anti-CTLA-4 combined therapy. Cancer Cell 35, 238&#x2013;255 (2019).\" href=\"#ref-CR90\" id=\"ref-link-section-d74592869e5970_1\">90<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Riaz, N. et al. Tumor and microenvironment evolution during immunotherapy with nivolumab. Cell 171, 934&#x2013;949 (2017).\" href=\"#ref-CR91\" id=\"ref-link-section-d74592869e5970_2\">91<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Lee, J. S. et al. Synthetic lethality-mediated precision oncology via the tumor transcriptome. Cell 184, 2487&#x2013;2502 (2021).\" href=\"#ref-CR92\" id=\"ref-link-section-d74592869e5970_3\">92<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Liu, D. et al. Integrative molecular and clinical modeling of clinical outcomes to PD1 blockade in patients with metastatic melanoma. Nat. Med. 25, 1916&#x2013;1927 (2019).\" href=\"#ref-CR93\" id=\"ref-link-section-d74592869e5970_4\">93<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Campbell, K. M. et al. Prior anti-CTLA-4 therapy impacts molecular characteristics associated with anti-PD-1 response in advanced melanoma. Cancer Cell 41, 791&#x2013;806 (2023).\" href=\"#ref-CR94\" id=\"ref-link-section-d74592869e5970_5\">94<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Mariathasan, S. et al. TGF&#x3B2; attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells. Nature 554, 544&#x2013;548 (2018).\" href=\"#ref-CR95\" id=\"ref-link-section-d74592869e5970_6\">95<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Cui, C. et al. Ratio of the interferon-&#x3B3; signature to the immunosuppression signature predicts anti-PD-1 therapy response in melanoma. npj Genom. Med. 6, 7 (2021).\" href=\"#ref-CR96\" id=\"ref-link-section-d74592869e5970_7\">96<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"He, Y. et al. Multi-omics characterization and therapeutic liability of ferroptosis in melanoma. Signal Transduct. Target. Ther. 7, 268 (2022).\" href=\"#ref-CR97\" id=\"ref-link-section-d74592869e5970_8\">97<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Kang, J. et al. Systematic dissection of tumor-normal single-cell ecosystems across a thousand tumors of 30 cancer types. Nat. Commun. 15, 4067 (2024).\" href=\"#ref-CR98\" id=\"ref-link-section-d74592869e5970_9\">98<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 99\" title=\"Ravi, A. et al. Genomic and transcriptomic analysis of checkpoint blockade response in advanced non-small cell lung cancer.Nat. Genet.55, 807&#x2013;819 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR99\" id=\"ref-link-section-d74592869e5973\" rel=\"nofollow noopener\" target=\"_blank\">99<\/a>, representing four cancer types (melanoma and three carcinoma types) (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>). All expression data were normalized to TPM before analysis.<\/p>\n<p>Using the Spatial EcoTyper deconvolution model, we predicted SE abundances across tumours in each dataset. We also evaluated the activity of publicly available transcriptional features associated with immunotherapy response<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 100\" title=\"Liu, Y. et al. Predicting patient outcomes after treatment with immune checkpoint blockade: a review of biomarkers derived from diverse data modalities. Cell Genom. 4, 100444 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR100\" id=\"ref-link-section-d74592869e5983\" rel=\"nofollow noopener\" target=\"_blank\">100<\/a>, including carcinoma ecotypes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e5987\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>, T-cell dysfunction<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Jiang, P. et al. Signatures of T cell dysfunction and exclusion predict cancer immunotherapy response. Nat. Med. 24, 1550&#x2013;1558 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR58\" id=\"ref-link-section-d74592869e5991\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>, T-cell exclusion, microsatellite instability (MSI)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Jiang, P. et al. Signatures of T cell dysfunction and exclusion predict cancer immunotherapy response. Nat. Med. 24, 1550&#x2013;1558 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR58\" id=\"ref-link-section-d74592869e5995\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>, tumour immune dysfunction and exclusion (TIDE)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Jiang, P. et al. Signatures of T cell dysfunction and exclusion predict cancer immunotherapy response. Nat. Med. 24, 1550&#x2013;1558 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR58\" id=\"ref-link-section-d74592869e5999\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>, immune resistance signatures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 101\" title=\"Jerby-Arnon, L. et al. A cancer cell program promotes T cell exclusion and resistance to checkpoint blockade. Cell 175, 984&#x2013;997 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR101\" id=\"ref-link-section-d74592869e6004\" rel=\"nofollow noopener\" target=\"_blank\">101<\/a>, IMPRES<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 102\" title=\"Auslander, N. et al. Robust prediction of response to immune checkpoint blockade therapy in metastatic melanoma. Nat. Med. 24, 1545&#x2013;1549 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR102\" id=\"ref-link-section-d74592869e6008\" rel=\"nofollow noopener\" target=\"_blank\">102<\/a>, TLS signatures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 57\" title=\"Cabrita, R. et al. Tertiary lymphoid structures improve immunotherapy and survival in melanoma. Nature 577, 561&#x2013;565 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR57\" id=\"ref-link-section-d74592869e6012\" rel=\"nofollow noopener\" target=\"_blank\">57<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 103\" title=\"Meylan, M. et al. Tertiary lymphoid structures generate and propagate anti-tumor antibody-producing plasma cells in renal cell cancer. Immunity 55, 527&#x2013;541 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR103\" id=\"ref-link-section-d74592869e6015\" rel=\"nofollow noopener\" target=\"_blank\">103<\/a>, cytolytic score (GZMA and PRF1)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 104\" title=\"Rooney, M. S., Shukla, S. A., Wu, C. J., Getz, G. &amp; Hacohen, N. Molecular and genetic properties of tumors associated with local immune cytolytic activity. Cell 160, 48&#x2013;61 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR104\" id=\"ref-link-section-d74592869e6025\" rel=\"nofollow noopener\" target=\"_blank\">104<\/a>, MHC-I signature (HLA-A, HLA-B, HLA-C, B2M and CASP8), PD-L1 (CD274), 18-gene inflammatory signatures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 55\" title=\"Ayers, M. et al. IFN-&#x3B3;-related mRNA profile predicts clinical response to PD-1 blockade. J. Clin. Invest. 127, 2930&#x2013;2940 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR55\" id=\"ref-link-section-d74592869e6049\" rel=\"nofollow noopener\" target=\"_blank\">55<\/a>, combined tumour and immune signals (MAP4K and TBX3)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 105\" title=\"Freeman, S. S. et al. Combined tumor and immune signals from genomes or transcriptomes predict outcomes of checkpoint inhibition in melanoma. Cell Rep. Med. 3, 100500 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR105\" id=\"ref-link-section-d74592869e6059\" rel=\"nofollow noopener\" target=\"_blank\">105<\/a>, an M1 macrophage signature<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 56\" title=\"Hwang, S. et al. Immune gene signatures for predicting durable clinical benefit of anti-PD-1 immunotherapy in patients with non-small cell lung cancer. Sci. Rep. 10, 643 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR56\" id=\"ref-link-section-d74592869e6063\" rel=\"nofollow noopener\" target=\"_blank\">56<\/a> and an IFNG signature<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 55\" title=\"Ayers, M. et al. IFN-&#x3B3;-related mRNA profile predicts clinical response to PD-1 blockade. J. Clin. Invest. 127, 2930&#x2013;2940 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR55\" id=\"ref-link-section-d74592869e6071\" rel=\"nofollow noopener\" target=\"_blank\">55<\/a>. The activity of carcinoma ecotypes from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Luca, B. A. et al. Atlas of clinically distinct cell states and ecosystems across human solid tumors. Cell 184, 5482&#x2013;5496 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR6\" id=\"ref-link-section-d74592869e6075\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>, T-cell dysfunction, exclusion, MSI and TIDE from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Jiang, P. et al. Signatures of T cell dysfunction and exclusion predict cancer immunotherapy response. Nat. Med. 24, 1550&#x2013;1558 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR58\" id=\"ref-link-section-d74592869e6079\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a>, immune resistance signatures from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 101\" title=\"Jerby-Arnon, L. et al. A cancer cell program promotes T cell exclusion and resistance to checkpoint blockade. Cell 175, 984&#x2013;997 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR101\" id=\"ref-link-section-d74592869e6083\" rel=\"nofollow noopener\" target=\"_blank\">101<\/a> and IMPRES from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 102\" title=\"Auslander, N. et al. Robust prediction of response to immune checkpoint blockade therapy in metastatic melanoma. Nat. Med. 24, 1545&#x2013;1549 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR102\" id=\"ref-link-section-d74592869e6087\" rel=\"nofollow noopener\" target=\"_blank\">102<\/a> was evaluated using their respective algorithms with default settings. For the remaining features, average gene expression was computed using log2 TPM data.<\/p>\n<p>The association between each feature and ICI response was assessed using a z-score derived from a two-sided Wilcoxon rank-sum test, within each dataset. These z-scores were then combined across datasets using Liptak\u2019s method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 106\" title=\"Lipt&#xE1;k, T. On the combination of independent tests. A Magyar Tudom&#xE1;nyos Akad&#xE9;mia Matematikai Kutat&#xF3; Int&#xE9;zet&#xE9;nek K&#xF6;zlem&#xE9;nyei 3, 171&#x2013;197 (1958).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR106\" id=\"ref-link-section-d74592869e6102\" rel=\"nofollow noopener\" target=\"_blank\">106<\/a>, weighted by the square root of sample sizes. The resulting combined z-scores were converted to two-sided P-values (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>). For the analysis in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3h<\/a>, data from all four cancer types were included, whereas the comparison in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5b<\/a> was restricted to melanoma datasets.<\/p>\n<p>SE levels versus TMB and CD274 expression in ICI-treated patients<\/p>\n<p>Of the pretreatment bulk RNA-seq tumour profiles analysed in \u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec71\" rel=\"nofollow noopener\" target=\"_blank\">Associations with immunotherapy response\u2019<\/a>, 465 patients from four studies with melanoma (n\u2009=\u2009150), non-small cell lung cancer (n\u2009=\u200943) or bladder cancer (n\u2009=\u2009272) have whole-exome-sequencing-derived TMB data available (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>). TMB values, defined as the total number of non-synonymous mutations per patient, and CD274 (encoding PD-L1) expression levels were log2-transformed. For each dataset, univariate Cox proportional hazards regression models were fitted to evaluate the association between standardized feature levels (SE7, SE8, SE4, TMB and CD274 expression) and overall survival. The resulting HRs and their associated standard errors were pooled across datasets within each cancer type, and across cancer types, using a nested random-effects meta-analysis implemented in the rma.mv function of the metafor R package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 107\" title=\"Viechtbauer, W. Conducting meta-analyses in R with the metafor package. J. Stat. Softw. 36, 1&#x2013;48 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR107\" id=\"ref-link-section-d74592869e6158\" rel=\"nofollow noopener\" target=\"_blank\">107<\/a> (v.4.8.0), with default parameters. Specifically, for each covariate, we used rma.mv to combine the log-hazard ratios and their corresponding variances (standard error squared) across outer and inner grouping factors (cancer type and dataset, respectively), then extracted pooled HRs, 95% confidence intervals and associated P-values to generate the forest plot in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3i<\/a>. Because all covariates were standardized, each HR reflects the association with overall survival for the same (1\u2009s.d.) change in the predictor, enabling direct comparison of their relative influence. For the analysis in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig14\" rel=\"nofollow noopener\" target=\"_blank\">9f<\/a>, multivariable Cox proportional HR models were applied to each dataset, including standardized levels of SE8, SE7 or SE4 jointly with TMB and CD274 expression. The resulting log-HRs and standard errors were combined using the same nested random-effects framework described above.<\/p>\n<p>Liquid EcoTyper framework<\/p>\n<p>Current analyses of the tumour microenvironment rely on invasive solid-tumour biopsies, which are prone to sampling bias and generally restricted to a single diagnostic biopsy. This limitation hinders the application of SEs as biomarkers in clinical settings. To address these challenges, we developed Liquid EcoTyper, a deep-learning framework for non-invasive profiling of SEs using plasma cfDNA methylation profiles.<\/p>\n<p>Liquid EcoTyper is built around a CpG set binary network (CSBN), in which informative CpG sets and associated weights are learnt simultaneously within a unified framework to enable multivariate prediction of SE levels (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>). This CSBN approach draws on the gene set binary network model originally introduced for predicting cellular developmental potential from single-cell RNA-seq data<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 108\" title=\"Kang, M. et al. Improved reconstruction of single-cell developmental potential with CytoTRACE 2. Nat. Methods 22, 2258&#x2013;2263 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR108\" id=\"ref-link-section-d74592869e6191\" rel=\"nofollow noopener\" target=\"_blank\">108<\/a>. Notably, sample methylation profiles are encoded at the CpG set level for model inference, analogous to gene sets in the previous study. This representation improves robustness to batch effects and technical dropout in methylation sequencing data, and enhances generalizability across data types, including both tumour and plasma methylation profiles.<\/p>\n<p>Liquid EcoTyper network architectureInput and output<\/p>\n<p>As input, the model takes an n\u2009\u00d7\u2009s matrix X containing the preprocessed methylation levels of n CpGs over s samples (section \u2018Methylation data preprocessing\u2019 below). At evaluation, the model yields an s\u2009\u00d7\u2009l matrix \u0176 containing the predicted levels of l classes over s samples. Model classes include SEs, non-SE (section \u2018Recovery of SE-specific cell states\u2019 above) and a \u2018background\u2019 class representing DNA not derived from the tumour compartment (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>).<\/p>\n<p>CpG set binary network<\/p>\n<p>Model input X is passed to a core binary module in which m CpG sets are learnt in binary n\u2009\u00d7\u2009m matrix WB. As described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 108\" title=\"Kang, M. et al. Improved reconstruction of single-cell developmental potential with CytoTRACE 2. Nat. Methods 22, 2258&#x2013;2263 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR108\" id=\"ref-link-section-d74592869e6268\" rel=\"nofollow noopener\" target=\"_blank\">108<\/a>, WB has a continuous equivalent W used during training for model initialization and back-propagation that undergoes binarization at each forward pass. Here, CpG sets encoded within WB are scored simply by averaging normalized methylation values over the selected CpGs per sample:<\/p>\n<p>$$S:={\\rm{s}}{\\rm{c}}{\\rm{o}}{\\rm{r}}{\\rm{e}}(X{,W}^{B})\\in {{\\mathbb{R}}}^{s\\times m},$$<\/p>\n<p>$${S}_{\\bullet ,1\\le j\\le m}=\\frac{{X}^{{\\rm{\\top }}}{W}_{\\bullet ,j}^{B}}{{\\parallel {W}_{\\bullet ,j}^{B}\\parallel }_{1}},$$<\/p>\n<p>where\\(\\,{S}_{\\bullet ,j}\\) and \\({W}_{\\bullet ,j}^{B}\\) denote the j-th columns of the respective matrices. Scores are standardized across samples by batch normalization, yielding s\u2009\u00d7\u2009m score matrix Snorm.<\/p>\n<p>Prediction layer<\/p>\n<p>The CpG set scores encoded in Snorm are then passed through a linear layer to produce s \u00d7 l matrix Q. This matrix is then transformed using the sigmoid function \u03c3 and rescaled to yield the final output prediction \u0176:<\/p>\n<p>$$P={\\rm{\\sigma }}(Q)+0.001\\in {{\\mathbb{R}}}^{s\\times l},{\\hat{Y}}_{1\\le i\\le s,\\bullet }=\\frac{{P}_{i,\\bullet }}{\\parallel {{P}_{i,\\bullet }\\parallel }_{1}}.$$<\/p>\n<p>where \\({\\hat{Y}}_{i,\\bullet }\\) and \\({P}_{i,\\bullet }\\) denote the i-th rows of the respective matrices.<\/p>\n<p>Liquid EcoTyper training<\/p>\n<p>The Liquid EcoTyper model was implemented and trained using PyTorch 2.2.0. In practice, the model feature space is of size n\u2009=\u200938,431 CpGs and the CpG set binary module included m\u2009=\u2009400 CpG sets.<\/p>\n<p>Loss function<\/p>\n<p>For model training we define a custom loss function incorporating both the mean cross-SE Pearson correlation per sample and the mean cross-sample Pearson correlation per SE. In combination with the model structure, this loss function prioritizes robust maintenance of linear relationships, particularly of SE levels across samples\u2014essential for clinical utility\u2014while discouraging overfitting to training values. In detail, we define prediction loss as follows:<\/p>\n<p>$$\\begin{array}{c}{\\rm{P}}{\\rm{r}}{\\rm{e}}{\\rm{d}}{\\rm{L}}{\\rm{o}}{\\rm{s}}{\\rm{s}}(\\hat{Y},Y)=1-\\frac{1}{s}\\mathop{\\sum }\\limits_{i=1}^{s}{\\rm{P}}{\\rm{e}}{\\rm{a}}{\\rm{r}}{\\rm{s}}{\\rm{o}}{\\rm{n}}({\\hat{Y}}_{i}^{{\\rm{\\top }}},{Y}_{i}^{{\\rm{\\top }}})\\\\ \\,\\,\\,\\,\\,\\,\\,\\,\\,+\\,1-\\frac{1}{l}\\mathop{\\sum }\\limits_{j=1}^{l}{\\rm{P}}{\\rm{e}}{\\rm{a}}{\\rm{r}}{\\rm{s}}{\\rm{o}}{\\rm{n}}({\\hat{Y}}_{j},{Y}_{j}),\\end{array}$$<\/p>\n<p>where \u0176 and Y denote predicted and true sample level matrices, respectively, and subscripts indicate matrix columns. Along with the prediction loss, we also include a term penalizing CpG set size as described previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"Oliveira, M. F. de et al. High-definition spatial transcriptomic profiling of immune cell populations in colorectal cancer. Nat. Genet. 57, 1512&#x2013;1523 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR78\" id=\"ref-link-section-d74592869e7189\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a>, designed to provide additional regularization to the model. Full model loss is then computed as<\/p>\n<p>$${\\rm{L}}{\\rm{o}}{\\rm{s}}{\\rm{s}}(\\hat{Y},Y,{W}^{B})={\\rm{P}}{\\rm{r}}{\\rm{e}}{\\rm{d}}{\\rm{L}}{\\rm{o}}{\\rm{s}}{\\rm{s}}(\\hat{Y},Y)+\\lambda {\\parallel \\frac{1}{n}{({W}^{B})}^{{\\rm{\\top }}}({W}^{B})\\odot I\\parallel }_{F},$$<\/p>\n<p>where \\(\\lambda \\) denotes the CpG set size penalty weight, I denotes the m\u2009\u00d7\u2009m identity matrix, \u2299 denotes the Hadamard product, and \\({\\parallel \\bullet \\parallel }_{F}\\) denotes the Frobenius norm. In practice, \\(\\lambda \\) was set to \\(\\sqrt{10}\\) to balance loss terms.<\/p>\n<p>Model regularization<\/p>\n<p>To support robustness to technical drop-out in EM-seq data, drop-out was applied to model input matrix X during training at a rate of 0.5.<\/p>\n<p>Model initialization and updates<\/p>\n<p>Model weights were given default initialization except binary weight matrix W, which here was initialized with values sampled from the Gaussian distribution with mean \u03bc\u2009=\u2009\u20130.125 and standard deviation \u03c3\u2009=\u20090.055 to give a highly sparse initial matrix targeting 400\u2013500 CpGs initially selected per set. At each iteration, model parameters were updated using PyTorch\u2019s NAdam optimizer with learning rate lr\u2009=\u20090.001 and cross-epoch gradient accumulation given the stabilizing role of inertia in training binary neural networks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 109\" title=\"Alizadeh, M., Fern&#xE1;ndez-Marqu&#xE9;s, J., Lane, N. D. &amp; Gal, Y. A systematic study of binary neural networks&#x2019; optimisation. University of Oxford &#010;                https:\/\/www.cs.ox.ac.uk\/publications\/publication13850-abstract.html&#010;                &#010;               (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR109\" id=\"ref-link-section-d74592869e7514\" rel=\"nofollow noopener\" target=\"_blank\">109<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 110\" title=\"Helwegen, K. et al. Latent weights do not exist: rethinking binarized neural network optimization. In Proc. Advances in Neural Information Processing Systems 32 (eds Wallach, H. et al.) (NeurIPS, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR110\" id=\"ref-link-section-d74592869e7517\" rel=\"nofollow noopener\" target=\"_blank\">110<\/a>.<\/p>\n<p>Model training, evaluation and stopping<\/p>\n<p>Models were trained over ten random splits (80% training, 20% validation) of the simulated training cohort (section \u2018Simulation of plasma cfDNA with tumour contribution\u2019 below). Each model was trained for 40 epochs, with model performance by epoch evaluated over training and validation sets using the PredLoss function described above (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec79\" rel=\"nofollow noopener\" target=\"_blank\">Loss function\u2019\u00a0above<\/a>). Performance on the validation split was used for early stopping. Final model weights were selected corresponding to the epoch yielding best validation performance.<\/p>\n<p>Model ensembling<\/p>\n<p>For all evaluations outside the training framework, the outputs of all ten models (from ten random folds) are averaged to yield a single ensembled prediction matrix.<\/p>\n<p>Initial feature selection<\/p>\n<p>To prepare a reduced model feature space for efficient training, informative CpGs per SE were selected from the TCGA melanoma cohort, which includes paired methylation and bulk RNA-seq profiles (n\u2009=\u2009461 tumours) obtained from the TCGA PanCanAtlas (<a href=\"https:\/\/gdc.cancer.gov\/about-data\/publications\/pancanatlas\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/gdc.cancer.gov\/about-data\/publications\/pancanatlas<\/a>). First, the set of CpGs was reduced to those detected across all TCGA melanoma tumours. For each SE, we then performed differential methylation analysis in the training cohort to identify differentially methylated CpGs. Specifically, for each SE class, we grouped samples with inferred SE abundance from paired bulk RNA-seq data above the 75th or below the 25th quantile. We then assessed each CpG for differential methylation between groups. We computed the significance of group difference by the Wilcoxon rank-sum test and the magnitude by absolute value of the difference of group means, \\(\\Delta =|{\\mu }_{75}-{\\mu }_{25}|\\), selecting CpGs satisfying P\u2009&lt;\u20090.05 and \\(\\Delta \\)\u2009&gt;\u20090.1. If more than 5,000 CpGs met both criteria for an SE, we selected the top 5,000 by differential methylation magnitude. Selection of up to 5,000 CpGs was done separately for positive and negative differential methylation, allowing for a total of up to 10,000 CpGs to be selected per SE. To ensure adequate features for SE recovery from methylation data, we required a minimum of 1,000 informative CpGs per SE. In practice, ecotype SE6 did not meet this criterion for the melanoma model and was omitted. In total, this yielded a feature space of 38,431 CpGs.<\/p>\n<p>Methylation data preprocessingMapping EM-seq CpGs to HM450K probe IDs<\/p>\n<p>Methylation observations from EM-seq data, given in terms of CpG locations (chromosome, start and end coordinates), were matched to HM450 probe IDs using the GRCh38 HM450 manifest downloaded from <a href=\"https:\/\/zwdzwd.github.io\/InfiniumAnnotation\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/zwdzwd.github.io\/InfiniumAnnotation<\/a>. In brief, the manifest was reduced to entries with defined values for CpG location as well as HM450 probe ID. Any duplicate entries per CpG location were subsequently dropped, yielding a one-to-one mapping.<\/p>\n<p>Normalization and imputation<\/p>\n<p>Data were subset to model features with methylation values subtracted from one and normalized to zero mean and unit variance over model CpGs per sample. Missing values were then imputed per CpG by dataset mean when defined and replaced by zeros otherwise.<\/p>\n<p>Simulation of plasma cfDNA with tumour contribution<\/p>\n<p>Because Liquid EcoTyper requires training data with known composition, and because of the limited availability of paired tumour and plasma cfDNA data, we simulated a training dataset by combining plasma cfDNA and tumour methylation profiles (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4b<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10a<\/a>). In doing so, we considered several factors. First, different malignant and TME cell states may exhibit variable kinetics of DNA shedding into circulation. We therefore prioritized cross-patient relative levels of SEs during training (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec79\" rel=\"nofollow noopener\" target=\"_blank\">Loss function\u2019\u00a0above<\/a>). Second, and relatedly, real biopsy samples, including those from patients\u00a0with metastatic cancer, are prone to sampling bias, so plasma and ground-truth tumour SEs need not match exactly. Third, even cfDNA samples from patients with cancer\u00a0with undetectable circulating tumour DNA by mutation analysis may include some tumour (and TME) methylation signal. Therefore, to isolate TME-derived signal during model training, we leveraged healthy donor plasma cfDNA to serve as a background class. Finally, inter-subject variation is much lower among cfDNA samples from healthy controls than from patients with cancer\u00a0in our data (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>). We therefore considered it justifiable to combine methylomes from different healthy individuals to increase variation and boost the size of the training cohort (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec89\" rel=\"nofollow noopener\" target=\"_blank\">Generation of additional background cfDNA profiles\u2019\u00a0below<\/a>).<\/p>\n<p>Accordingly, we designed simulated mixtures to contain both \u2018background\u2019 cfDNA and tumour tissue contributions, with the background compartment derived from plasma cfDNA profiled from cohort 2 healthy individuals (n\u2009=\u200923 in total; partitioned into n\u2009=\u200918 for training cohort and n\u2009=\u20095 for test cohort) and the tumour compartment derived from the TCGA melanoma methylation cohort (n\u2009=\u2009461 in total; partitioned into n\u2009=\u2009346 for training cohort and n\u2009=\u2009115 for test cohort) (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4b,c<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10a<\/a>). Paired bulk RNA-seq data from TCGA tumour samples enabled the recovery of tissue SE levels through bulk deconvolution, as described in section\u00a0\u2018NMF model training for bulk deconvolution\u2019\u00a0above.<\/p>\n<p>Generation of additional background cfDNA profiles<\/p>\n<p>Given the rationale outlined above, to expand the pool of healthy cfDNA profiles, we prepared new profiles from the samples of each cohort separately to reach the same number of samples as available from tumour tissue. For each new profile, we randomly selected three samples from the appropriate cohort (training or test), perturbing the methylation profiles of each sample with noise to reduce collinearity across new mixtures, then combining according to randomly generated fractions.<\/p>\n<p>In detail, fractional composition was generated by sampling from the unit interval uniformly at random for each sample, then rescaling the three values to sum to one, yielding mixing fractions \u03d51, \u03d52 and \u03d53. Noise perturbations were applied multiplicatively to raw methylation values, with noise level parameters \u03bd1, \u03bd2 and \u03bd3 generated by sampling from the normal distribution. Given sample methylation profiles M1, M2 and M3, perturbed profiles \\({\\hat{M}}_{1},{\\hat{M}}_{2}\\,{\\rm{and}}\\,{\\hat{M}}_{3}\\) were computed as follows:<\/p>\n<p>$${\\hat{M}}_{i}=min((1+s{\\nu }_{i}){M}_{i},1),$$<\/p>\n<p>where s denotes a scaling parameter, set here to s\u2009=\u20090.02, and min operates element-wise to cap resulting methylation values at 1. With these mixing parameters defined, the resulting new healthy cfDNA profile Mnew is given by:<\/p>\n<p>$${M}_{{\\rm{n}}{\\rm{e}}{\\rm{w}}}=\\mathop{\\sum }\\limits_{i=1}^{3}{\\phi }_{i}{\\hat{M}}_{i}.$$<\/p>\n<p>Combination of background and tumour compartments<\/p>\n<p>With the resulting background cfDNA profiles now matching the tumour tissue profiles in number, we generated the final sets of simulated samples by matching background and tumour profiles one-to-one in the training and test cohorts (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10a<\/a>). For each simulated sample, we generate mixing fractions \u03d5T and \u03d5B of tumour and background, respectively, by sampling \u03d5B uniformly at random from a desired range, then computing \u03d5T\u2009=\u20091\u2009\u2212\u2009\u03d5B as the complement. This range was selected here as [0.2, 0.6] to weight fractions in favour of tumour contribution while preserving substantial background for regularization to support model extensibility to both tumour tissue and plasma cfDNA methylation profiles. Of note, tumour samples contain highly variable malignant cell content (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec92\" rel=\"nofollow noopener\" target=\"_blank\">Validation of simulated samples\u2019<\/a> below), capturing a broad range of both malignant and TME DNA (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>). As such, this approach enables sufficient exposure of the model to diverse tumour fractions during training, facilitating generalizability. The resulting simulated samples are then constructed as:<\/p>\n<p>$${M}_{{\\rm{n}}{\\rm{e}}{\\rm{w}}}={\\phi }_{T}{M}_{T}+{\\phi }_{B}{M}_{B},$$<\/p>\n<p>where MT denotes the raw methylation profile of the selected tumour tissue sample and MB denotes the raw methylation profile of the selected background profile, newly generated as described above.<\/p>\n<p>Ground-truth composition of simulated samples<\/p>\n<p>Tumour-specific SE levels used for simulating cfDNA above were inferred by applying the Spatial EcoTyper deconvolution model to tumour RNA-seq profiles as described above\u00a0in section\u00a0\u2018NMF model training for bulk deconvolution\u2019. Following this, any SE with inadequate feature support in Liquid EcoTyper (for example, SE6 in the melanoma-specific model) was omitted and the remaining sample-level deconvolution results were rescaled to unit sum. To define ground-truth SE levels in simulated plasma cfDNA samples, we then scaled the SE levels in corresponding tumour samples by \u03d5T, with the ground-truth background level given by \u03d5B, for each simulated sample.<\/p>\n<p>Validation of simulated samples<\/p>\n<p>To determine whether simulated profiles are a suitable proxy for authentic cfDNA profiles, we performed a direct assessment against real cfDNA samples from healthy individuals and patients\u00a0with melanoma. Profiles from cohort 2 healthy individuals (n\u2009=\u200923), cohort 3 patients\u00a0with melanoma (n\u2009=\u200978) and simulated training and test cohorts (n\u2009=\u2009461) were jointly embedded by PCA. This embedding revealed tight clustering of healthy plasma cfDNA methylomes, with some melanoma cfDNA samples overlapping the healthy region but many others scattering out further (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>, left). Simulated cfDNA methylomes overlapped real melanoma cfDNA methylomes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>, left).<\/p>\n<p>Given these results, we hypothesized that the embedding distribution was organized by sample tumour content. Thus, we extended the visualization to colour by ctDNA percentage (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>, right). For real cfDNA samples, ctDNA percentage was given by AVENIO for patients with melanoma\u00a0(n\u2009=\u200960; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>) and set to zero for healthy individuals. We computed an effective ctDNA percentage for simulated profiles by multiplying the total tumour fraction by the tumour purity using consensus measurement of purity estimations<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 111\" title=\"Aran, D., Sirota, M. &amp; Butte, A. J. Systematic pan-cancer analysis of tumour purity. Nat. Commun. 6, 8971 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR111\" id=\"ref-link-section-d74592869e8269\" rel=\"nofollow noopener\" target=\"_blank\">111<\/a>, available for 456 of 461 samples. The resulting embedding revealed a gradient of tumour content across the embedding for both real and simulated melanoma cfDNA samples, with samples harbouring the lowest tumour content placed closer to healthy plasma (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10b<\/a>, right).<\/p>\n<p>Performance assessment of Liquid EcoTyperApplication to held-out simulated data<\/p>\n<p>We tested Liquid EcoTyper\u2019s ability to recover ground-truth levels of SEs from plasma methylation profiles by application to the held-out test cohort of simulated data (n\u2009=\u2009115; section\u00a0\u2018Simulation of plasma cfDNA with tumour contribution\u2019\u00a0above). Performance for each SE individually, as well as for recovery of the underlying cross-correlation structure of SE levels, was assessed by Spearman correlation between ground-truth and predicted levels (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4c<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10c<\/a>).<\/p>\n<p>Extension to carcinomas<\/p>\n<p>To evaluate generalizability, we prepared Liquid EcoTyper models for 13 carcinoma types profiled by TCGA, each with paired bulk RNA-seq and methylation arrays for more than 100 tumour samples. Colon adenocarcinoma (COAD) and rectum adenocarcinoma (READ) were combined as colorectal cancer (CRC). For each carcinoma type, we repeated the procedures described in\u00a0sections \u2018Initial feature selection\u2019, \u2018Simulation of plasma cfDNA with tumour contribution\u2019, \u2018Model training\u2019 and \u2018Application to held-out simulated data\u2019 above. As detailed in \u2018Initial feature selection\u2019, any SE with initial feature set selection resulting in fewer than 1,000 associated CpGs was excluded from model training and evaluation (SE1 from two carcinomas, SE3 from one carcinoma, SE6 from four and SE9 from one). We aggregated model performance results across carcinoma types by average Spearman correlation coefficients between predicted (Liquid EcoTyper) and expected SE levels for held-out test sets (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10d<\/a>, left).<\/p>\n<p>We also evaluated a pan-carcinoma leave-one-out framework, training Liquid EcoTyper models over 12 carcinoma types at a time (150 randomly selected tumour samples per type to balance representation, totalling 1,800 training samples) and evaluating performance on each held-out carcinoma type in turn. For each pan-carcinoma model, we followed the same process as described above, with the exception that the initial CpG feature set for each cohort was selected according to a consensus mechanism across included carcinoma types. For each SE and methylation direction, CpGs were ranked by selection frequency across per-carcinoma feature sets and the top 5,000 were included. With this approach, all SEs had adequate CpG feature set coverage. To promote a fair comparison with carcinoma-specific models (above), for each held-out carcinoma type c, we excluded any SEs from model evaluation that were also excluded from the model exclusively trained on carcinoma c. Results were aggregated as described above (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10d<\/a>, centre).<\/p>\n<p>Paired tumour and plasma from patients\u00a0with melanoma<\/p>\n<p>We assessed concordance of Liquid EcoTyper predictions for plasma cfDNA collected from 23 patients with melanoma\u00a0(cohort 2) (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4d<\/a>) against: first, SE levels inferred from matched tumour Visium or Visium HD by Spatial EcoTyper (15 patients); second, Liquid EcoTyper predictions from matched tumour EM-seq (20 patients); and third, Liquid EcoTyper predictions from matched PBMC EM-seq (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>), as shown in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4e\u2013i<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11b<\/a>.<\/p>\n<p>For five tumour samples, EM-seq was performed on both FFPE and fresh-frozen sections (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>). Given that SE levels inferred by Liquid EcoTyper were largely concordant across replicates (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11a<\/a>), we averaged SE levels across replicates for subsequent analyses. SE levels for Visium data were inferred as described above in \u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec67\" rel=\"nofollow noopener\" target=\"_blank\">Paired bulk RNA-seq and Visium ST\u2019<\/a> following quality control as described in \u201810x Visium (standard)\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>. For cases in which Visium replicates of adjacent sections were available (two patients, each with two samples profiled by Visium), we averaged SE levels across replicates (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>). Two Visium HD samples were also included in the comparison with paired plasma EM-seq data, subject to quality control, as described in \u201810x Visium HD data\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>. To avoid platform-specific variation in analysing Visium and Visium HD data jointly, SE abundances for Visium HD were inferred for 16-\u00b5m bins using the same Spatial EcoTyper deconvolution protocol as described for Visium (section\u00a0\u2018Paired bulk RNA-seq and Visium ST\u2019 above).<\/p>\n<p>Liquid EcoTyper outputs were renormalized to unit sum excluding inferred plasma cfDNA background levels. We then evaluated the consistency of SE levels between tumour ST and plasma cfDNA EM-seq (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4e\u2013h<\/a>), as well as between tumour EM-seq and plasma cfDNA EM-seq (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4e,i<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11b<\/a>), and between PBMC EM-seq and both tumour and plasma EM-seq (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4i<\/a>). To emphasize relative ordering, SE levels were ranked and normalized to the 0\u20131 range in each compartment, with concordance then quantified by Spearman correlation for each SE. All of the SE levels are provided in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>.<\/p>\n<p>To visualize the most prominent SE signals in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4h<\/a> (centre), the levels of each SE (SE7 or SE4) in the Visium data were first separately standardized to mean zero and unit variance across all spots and samples. Next, for each sample, the 95th percentile of SE7 and SE4 levels together was determined, and SE levels were binarized by setting all values greater than this threshold to one and all other values to zero. The bar plots in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4h<\/a> (bottom) show the means of the binarized data.<\/p>\n<p>Validation of Liquid EcoTyper through learnt CpG sets<\/p>\n<p>To further assess the ability of Liquid EcoTyper to learn biologically grounded methylation profiles of spatial cellular ecosystems, we performed a series of experiments leveraging the CpG sets learnt by the model to determine whether model predictions successfully target each SE accurately, specifically and in accordance with known biological features.<\/p>\n<p>Concordance with ground-truth associations<\/p>\n<p>We first assessed whether Liquid EcoTyper successfully learns and preserves the associations of each CpG set with ground-truth SE composition. For each CpG set extracted from the model, we averaged the methylation levels of its component CpGs in each training cohort sample, then compared the resulting quantities to ground-truth SE levels, for each SE, by Pearson correlation. We then replaced ground-truth with predicted SE levels and recomputed the associations for each CpG set\u2013SE pair. Concordance between associations with ground-truth and predicted SE levels was evaluated using Pearson correlation across all CpG sets, separately for each SE (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10e<\/a>).<\/p>\n<p>Model specificity<\/p>\n<p>To assess the ability of Liquid EcoTyper to target each SE specifically for predictions, we evaluated the impact of ablating the top associated CpG sets per SE on model performance. For each SE, we selected all CpG sets with ground-truth Pearson correlations greater than 0.25 in absolute value among the training cohort (as determined from \u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec98\" rel=\"nofollow noopener\" target=\"_blank\">Concordance with ground-truth associations\u2019<\/a> above) and set the corresponding entries of the binary matrix WB encoding the learnt CpG sets to zero for each ensemble model. We then applied the resulting model to the held-out test cohort methylomes, assessing Liquid EcoTyper specificity for the SE in terms of performance loss. We defined a performance loss index as the fractional difference in Spearman correlations \u03c1 of predicted versus ground-truth SE levels between the original (unablated) and ablated Liquid EcoTyper models:<\/p>\n<p>$$\\frac{{\\rho }_{{\\rm{o}}{\\rm{r}}{\\rm{i}}{\\rm{g}}{\\rm{i}}{\\rm{n}}{\\rm{a}}{\\rm{l}}}-{\\rho }_{{\\rm{a}}{\\rm{b}}{\\rm{l}}{\\rm{a}}{\\rm{t}}{\\rm{e}}{\\rm{d}}}}{{\\rho }_{{\\rm{o}}{\\rm{r}}{\\rm{i}}{\\rm{g}}{\\rm{i}}{\\rm{n}}{\\rm{a}}{\\rm{l}}}}.$$<\/p>\n<p>Performance loss was calculated for all SEs following the ablation of each SE individually and comparing the loss for the ablated SE (on-target) with that for other SEs (off-target) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10f<\/a>). For ease of interpretability, performance loss index values were capped to a range of 0 to 1, defined as performance losses of 0% (\u03c1ablated\u2009\u2265\u2009\u03c1original) and 100% (\u03c1ablated\u2009\u2264\u20090), respectively.<\/p>\n<p>Biological interpretability<\/p>\n<p>Hypomethylated CpGs proximal to the transcription start site (TSS) (for example, promoters, the first exon and the first intron) are generally associated with elevated gene expression<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 112\" title=\"Feinberg, A. P. &amp; Vogelstein, B. Hypomethylation distinguishes genes of some human cancers from their normal counterparts. Nature 301, 89&#x2013;92 (1983).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR112\" id=\"ref-link-section-d74592869e8599\" rel=\"nofollow noopener\" target=\"_blank\">112<\/a>. Because Liquid EcoTyper learns CpG sets that include promoters and gene bodies, and because the Illumina Infinium HumanMethylation450 (HM450) BeadChip array is enriched for regulatory and TSS-proximal CpGs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 113\" title=\"Bibikova, M. et al. High density DNA methylation array with single CpG site resolution. Genomics 98, 288&#x2013;295 (2011).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR113\" id=\"ref-link-section-d74592869e8603\" rel=\"nofollow noopener\" target=\"_blank\">113<\/a>, we proposed that, on average, learnt CpGs within these regions would reflect this general relationship. Therefore, to investigate links between CpG methylation and SE classes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10g\u2013i<\/a>), we focused on model CpGs (section \u2018Initial feature selection\u2019\u00a0above) that overlap with promoter regions, defined as the 1-kilobase region upstream of the TSS, and the gene body, with hg38 coordinates obtained from the HGNC track at <a href=\"https:\/\/genome.ucsc.edu\/cgi-bin\/hgTables\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/genome.ucsc.edu\/cgi-bin\/hgTables<\/a>.<\/p>\n<p>For each SE, we constructed a ranked gene list from model CpGs and ground-truth CpG set associations among the training cohort. To do so, each CpG set was first converted to a gene set comprising all of the unique HUGO gene symbols overlapping the selected CpGs in either promoter regions or gene body at least once. For each gene g within a given SE, we then computed the average ground-truth SE association (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec98\" rel=\"nofollow noopener\" target=\"_blank\">Concordance with ground-truth associations\u2019\u00a0above<\/a>) over all CpG sets in which gene g is represented, resulting in gene-level methylation scores for each SE. The gene-level scores for each SE\u2014along with non-SE and background components\u2014were converted into rank space with higher ranks denoting lower scores. This yielded a rank matrix R consisting of 8,860 evaluable genes (rows) by 10 components (columns). To prioritize SE-specific features, we then calculated, for each gene in R, the delta in rank space between each SE and the maximum rank among other components (the remaining SEs, non-SE and background). This resulted in a new matrix D, comprising deltas for each gene for all ten components. We then removed non-SE and background components from R and D, averaged the resulting matrices to exploit complementary information (ensuring that SEs of the same class were averaged), and median-subtracted each column to yield equal numbers of positive and negative ranks for each SE. Finally, we subjected the resulting SE-specific ranked lists to fgsea (v.1.25.1)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 114\" title=\"Korotkevich, G. et al. Fast gene set enrichment analysis. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/060012&#010;                &#010;               (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR114\" id=\"ref-link-section-d74592869e8645\" rel=\"nofollow noopener\" target=\"_blank\">114<\/a>, testing the enrichments of SE-specific consensus genes (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10g,h<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>). Global significance (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig15\" rel=\"nofollow noopener\" target=\"_blank\">10i<\/a>) was determined by calculating the number of top-ranked diagonal matches (row or column) across all possible permutations of the ranked lists, yielding an exact permutation P-value.<\/p>\n<p>Consistency of associations<\/p>\n<p>To assess faithfulness of model predictions to the learnt CpG set associations, we applied the procedure outlined above\u00a0in \u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec98\" rel=\"nofollow noopener\" target=\"_blank\">Concordance with ground-truth associations\u2019<\/a> to compute CpG set associations with predicted SE levels for melanoma plasma cfDNA collected as part of cohorts 2, 3, and 4, split by institution (in total, 79 samples collected at Yale and 17 at Washington University\u00a0in St. Louis; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>). We then compared the resulting CpG associations between institutions using Pearson correlation for each SE (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>).<\/p>\n<p>Robustness of Liquid EcoTyper<\/p>\n<p>To evaluate the robustness of plasma SE recovery to technical variation, we performed the following two experiments.<\/p>\n<p>Plasma SE recovery versus sequencing depth<\/p>\n<p>For the analysis in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11c<\/a>, we performed a down-sampling experiment using cohort 2 plasma samples with paired tumour ST profiling. Aligned read pairs were randomly downsampled using samtools<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 115\" title=\"Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR115\" id=\"ref-link-section-d74592869e8699\" rel=\"nofollow noopener\" target=\"_blank\">115<\/a> (v.1.18) at predefined fractions to achieve target sequencing depths of 20\u00d7, 15\u00d7, 10\u00d7 and 5\u00d7 for each sample. Down-sampled BAM files were subsequently processed through the remainder of the pipeline described in \u2018Enzymatic methyl sequencing\u2019 in\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>. Samples with an original sequencing depth below a given target were retained at their original depth. Liquid EcoTyper performance was then re-evaluated as described above\u00a0in \u2018Paired tumour and plasma from patients\u00a0with melanoma\u2019\u00a0using down-sampled plasma samples. Performance remained robust down to 10\u00d7 depth (the minimum depth in this study), with a significant but modest decline observed at 5\u00d7 depth (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11c<\/a>).<\/p>\n<p>Plasma SE recovery versus CpG imputation approach<\/p>\n<p>To evaluate whether the approach for imputing missing CpG methylation values (\u2018<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Sec85\" rel=\"nofollow noopener\" target=\"_blank\">Methylation data preprocessing\u2019\u00a0above<\/a>) influences Liquid EcoTyper predictions (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">11d,e<\/a>), we tested an alternative imputation strategy in which missing values were imputed uniformly at random from the range [0,1] before data normalization. Although the default approach leverages other methylation profiles in the cohort to infer missing values, this alternative uses no prior knowledge and instead adds random noise. Applied to cohort 2 plasma samples with paired tumour ST profiling, the random-imputation approach had only a small effect on samples with low coverage (higher imputation fraction), and there was no significant global difference in SE recovery performance, emphasizing robustness.<\/p>\n<p>Association between liquid SE levels and clinical outcomes in ICI-treated patients<\/p>\n<p>To evaluate the clinical relevance of SE levels in plasma cfDNA from patients with melanoma, we applied Liquid EcoTyper as described above\u00a0in \u2018Paired tumour and plasma from patients\u00a0with melanoma\u2019 to infer SE levels in 78 pretreatment plasma samples from cohort 3 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5a<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>) and 10 plasma samples from cohort 4 (Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>).<\/p>\n<p>Liquid SEs versus ICI response<\/p>\n<p>The association between each SE and ICI response was assessed using a two-sided Wilcoxon rank-sum test, comparing patients with DCB to those with NDB in each dataset (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5c,e<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig17\" rel=\"nofollow noopener\" target=\"_blank\">12a,g<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>, <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>) and by AUC (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig17\" rel=\"nofollow noopener\" target=\"_blank\">12f<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>).<\/p>\n<p>Liquid SEs versus survival<\/p>\n<p>For Kaplan\u2013Meier analyses (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5d,f<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig17\" rel=\"nofollow noopener\" target=\"_blank\">12b\u2013d<\/a>), we dichotomized patients based on the median level of each liquid SE (SE7, SE8 and SE4) in cohort 3 (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>). Differences in OS and PFS were assessed using a two-sided log-rank test with the survival (v.3.6.4) R package. Associations with OS and PFS were also evaluated using both univariate and multivariable Cox regression models, including models combining continuous liquid SE7, SE8 or SE4 levels with other clinical indices, such as sex, age, ICI type, melanoma subtype and BRAF mutation status. Results are provided in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig17\" rel=\"nofollow noopener\" target=\"_blank\">12e<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>. Additional Cox models evaluating the associations between OS and continuous liquid SE levels, ctDNA levels, TMB (log2-transformed number of non-synonymous mutations per megabase) and PD-L1 percentages (tumour proportion score, TPS), both alone and in bivariable models comparing liquid SE7, SE8 or SE4 levels against each of the above covariates, are provided in Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#Fig17\" rel=\"nofollow noopener\" target=\"_blank\">12h<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>. For PD-L1 TPS values reported as ranges (for example, 1\u20135), the median of the range was used. For PD-L1 TPS values reported as inequalities, less than x or more than x, we subtracted or added 0.5 to x, respectively (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a>). Finally, we performed a time-dependent AUC(t) analysis for comparative OS prediction using the timeROC function of the timeROC R package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 116\" title=\"Blanche, P., Dartigues, J.-F. &amp; Jacqmin-Gadda, H. Estimating and comparing time-dependent areas under receiver operating characteristic curves for censored event times with competing risks. Stat. Med. 32, 5381&#x2013;5397 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#ref-CR116\" id=\"ref-link-section-d74592869e8824\" rel=\"nofollow noopener\" target=\"_blank\">116<\/a> (v.0.4 with default parameters) (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>). The 6\u201318-month interval after ICI initiation was selected because it captures the minimum time needed for gauging durable clinical benefit (6 months) and provides additional time for delayed benefit. It also avoids losing patients to follow-up and minimizes the sparsity of events in cohort 3.<\/p>\n<p>Statistics and reproducibility<\/p>\n<p>Unless otherwise noted, all statistical tests were two-sided. Two-group comparisons were conducted with Wilcoxon rank-sum or signed-rank\u00a0tests as appropriate. For comparisons of SE levels imputed from tissue expression data (bulk, Visium ST and scRNA-seq), which are expected to be linearly related, Pearson correlations were applied to assess concordance. For tissue versus plasma comparisons of SE levels, Spearman correlations were used to assess the directional concordance of SE levels because strict assumptions of normality and linearity could not be guaranteed. For Cox regression models, the proportional hazards assumption was verified for each covariate by evaluating the Schoenfeld residuals.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10452-4#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Human subjects All human samples included in this study were collected with informed consent for research use and&hellip;\n","protected":false},"author":2,"featured_media":627883,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[24527,66026,9171,97,1159,30319,1877,1160,79],"class_list":["post-627882","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-biomarkers","tag-cancer-microenvironment","tag-genomics","tag-health","tag-humanities-and-social-sciences","tag-immunotherapy","tag-machine-learning","tag-multidisciplinary","tag-science"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/627882","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=627882"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/627882\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/627883"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=627882"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=627882"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=627882"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}