{"id":246737,"date":"2025-11-06T05:57:22","date_gmt":"2025-11-06T05:57:22","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/246737\/"},"modified":"2025-11-06T05:57:22","modified_gmt":"2025-11-06T05:57:22","slug":"epigenetically-driven-and-early-immune-evasion-in-colorectal-cancer-evolution","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/246737\/","title":{"rendered":"Epigenetically driven and early immune evasion in colorectal cancer evolution"},"content":{"rendered":"<p>Sample collection and sequencing<\/p>\n<p>Our FF\u2013WGS samples were comprised of processed data from previous sequencing experiments of our evolutionary predictions in CRC cohort, which has been described in refs. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2052\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Househam, J. et al. Phenotypic plasticity and genetic control in colorectal cancer evolution. Nature 611, 744&#x2013;753 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR23\" id=\"ref-link-section-d101659981e2055\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>. All patients gave informed consent in writing for collection of their materials to the UCLH Cancer Biobank (Research Ethics Committee approval 15\/YH\/0311). All investigators were blinded to patient data related to outcome, and all clinicopathological information.<\/p>\n<p>Our FFPE\u2013PS samples originated from 11 stage\u2009III microsatellite-stable cancers with lymph node metastases from the same cohort, 9 of which are shared with FF\u2013WGS samples. FFPE sections were cut in the following order: H&amp;E-1 (5\u2009\u03bcm), 5\u2009\u00d7\u20098\u2009\u03bcm for laser capture microdissection, H&amp;E-2 (5\u2009\u03bcm), 8\u2009\u00d7\u20095\u2009\u03bcm for CyCIF, H&amp;E-3 (5\u2009\u03bcm). The H&amp;E slides from each FFPE block were digitized using the NanoZoomer S210 or S60 (Hamamatsu). Images were reviewed using NDPViewer software (v.2.9.29).<\/p>\n<p>Regions of interest (ROIs) were identified as single glands or clusters of small adjacent glands (microbiopsies) from superficial, invasive margin or lymph node deposits. Superficial regions were defined as cancer regions adjacent to or contiguous with normal mucosa. The tumor\u2013normal interface was identified where possible, and the invasive margin was defined as the region within 500\u2009\u03bcm either side of the tumor\u2013normal interface (with an overall extent of ~1\u2009mm).<\/p>\n<p>Panel sequencing<\/p>\n<p>ROIs were microdissected using PALM MicroBeam Laser Microdissection (Zeiss). DNA was extracted using the High Pure FFPET DNA Isolation Kit (Roche). Extracted FFPE DNA was repaired using the NEBNext FFPE DNA Repair Mix. Postrepair, whole-genome libraries were prepared using the NEBNext Ultra II FS DNA Library Prep Kit for Illumina with unique molecular identifier (UMI) adapters ligated onto DNA molecules.<\/p>\n<p>PS on FFPE samples was carried out using a custom targeted panel designed by L. Zapata and manufactured by Twist BioSciences, focusing on regions encoding the immunopeptidome and antigen-presenting related genes. The immunopeptidome was defined as the set of the human nine-mers that bind strongly to one of the top 70 HLA alleles, was confirmed T\u2009cell positive in IEDB and were derived from a gene with mean expression of more than one fragments per kilobase million pancancer. The final list of these immunopeptidome loci can be obtained from: <a href=\"https:\/\/github.com\/luisgls\/SOPRANO\/blob\/master\/immunopeptidomes\/human\/allhlaBinders_exprmean1.IEDBpeps.unique.bed\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/luisgls\/SOPRANO\/blob\/master\/immunopeptidomes\/human\/allhlaBinders_exprmean1.IEDBpeps.unique.bed<\/a>.<\/p>\n<p>The Twist Target Enrichment Protocol was used to hybridize probes from the this custom panel with prepared libraries. Hybridized targets were isolated and amplified with PCR. Paired-end 50-bp runs were performed on Novaseq S1 (Illumina).<\/p>\n<p>Processing of sequencing data<\/p>\n<p>The processing of FF\u2013WGS data was detailed in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2097\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>.<\/p>\n<p>For FFPE\u2013PS samples, three sets of fastq reads were aligned to human reference genome build hg38 to generate an unmapped bam using FastqToBam (Fgbio v.1.3.0). Fastq files were then created from this unmapped bam (SamToFastq, Picard v.2.20.3) and aligned to human reference genome build GRCh38\/hg38 with Burrows\u2013Wheeler Aligner package BWA-MEM v.0.7.17. Alignment data from the outputted bam from this step were merged with data in the previously generated unmapped bam (using the MergeBamAlignment function of Picard v.2.20.3). The merged bam file was then used as input for GroupReadsByUMI (Fgbio v.1.3.0). \u2018Adjacency\u2019 was used as the strategy for grouping so that errors were allowed between UMIs, but only when there was a count gradient. The allowable number of edits\/changes between UMIs was set to 1, and minimum mapping quality to 30 (default).<\/p>\n<p>Consensus sequences were called from reads with the same unique molecular tag (CallMolecularConsensusReads) and filtered using FilterConsensusReads to exclude consensus sequences with fewer than two contributing reads, mask consensus bases with quality less than 30, accept a maximum raw-read error rate across the entire consensus read of 0.05, accept a maximum error rate for a single consensus base of 0.1 and accept a maximum fraction of 0.2 for no-calls in the read after filtering.<\/p>\n<p>The filtered reads were then re-aligned to GRCh38\/hg38, and variant calling was performed using Mutect2 (v.4.1.4.1) with a bed file specifying regions of the genome covered by the targeted panel. Normal samples from adjacent muscle were used as matched normal for variant calling. The resulting vcf files were merged for each patient and passed to Platypus (v.0.8.1.1) for multiregion variant calling.<\/p>\n<p>Variant calls were filtered to retain mutations within certain filtering criteria and adequate support for the variant, as described in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2114\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. The same filters were used for both FF\u2013WGS and FFPE\u2013PS samples, except for requiring at least five and at least eight reads covering each site in FF\u2013WGS and FFPE\u2013PS samples, respectively.<\/p>\n<p>FF\u2013WGS samples sequenced at shallow depth were genotyped and fitted on phylogenetic trees using the method described in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2121\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>.<\/p>\n<p>Raw RNA-seq read counts were normalized and converted using DESeq2 as described in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2128\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. The code for reproducing RNA-seq processing can be found at <a href=\"https:\/\/github.com\/JacobHouseham\/EPICC_transcriptome\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/JacobHouseham\/EPICC_transcriptome<\/a>. Raw counts were converted either to transcript per million (TPM) values or processed counts following VST, see \u20182.gene_expression_normalisation_and_filtering.Rmd\u2019 within the above repository. TPM values were used to assess expression levels of genes (for example, to establish constitutively expressed genes), while VST values were used for between-sample comparison of a particular gene.<\/p>\n<p>HLA haplotyping<\/p>\n<p>HLA-A, -B and -C haplotyping was performed using polysolver<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Shukla, S. A. et al. Comprehensive analysis of cancer-associated somatic mutations in class I HLA genes. Nat. Biotechnol. 33, 1152&#x2013;1158 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR47\" id=\"ref-link-section-d101659981e2155\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> (v.1.0) on FF\u2013WGS samples, by running shell_call_hla_type with default settings. As ethnicity information was not available, we used \u2018Unknown\u2019 for all samples. To increase coverage, we used merged bam files created from all sequencing files from a given cancer. For validation, we also performed haplotyping on merged bams formed of normal (blood or normal colon tissue) samples and compared the predicted haplotypes. The predicted haplotypes had a high concordance, with haplotypes predicted using all samples providing one more heterozygous haplotype than normal-only haplotypes in 6 of 30 cases. Based on the average homozygosity across CRCs (as seen in TCGA CRC samples<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Lakatos, E. et al. Evolutionary dynamics of neoantigens in growing tumors. Nat. Genet. 52, 1057&#x2013;1066 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR8\" id=\"ref-link-section-d101659981e2159\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>), we accepted the more heterozygous set of alleles predicted to define their set of HLA alleles. For all HLA alterations, we confirmed that the alterations are called independent of which haplotyping calls were used.<\/p>\n<p>For FFPE\u2013PS samples we used the calls derived from matched FF\u2013WGS samples. For the two patients where this was not available, we performed haplotyping on adjacent normal mucosal samples following the same steps as described above.<\/p>\n<p>Immune escape predictionMutations in HLA<\/p>\n<p>Somatic mutations in the HLA locus were predicted using polysolver<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Shukla, S. A. et al. Comprehensive analysis of cancer-associated somatic mutations in class I HLA genes. Nat. Biotechnol. 33, 1152&#x2013;1158 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR47\" id=\"ref-link-section-d101659981e2188\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> (v.1.0). The mutation detection script of polysolver (shell_call_hla_mutations_from_type) was run on matched tumor\u2013normal pairs to call tumor-specific alterations in HLA-aligned sequencing reads using MuTect (v.1.16). In addition, Strelka2<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 48\" title=\"Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat. Methods 15, 591&#x2013;594 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR48\" id=\"ref-link-section-d101659981e2192\" rel=\"nofollow noopener\" target=\"_blank\">48<\/a> (v.2.9.10) was run independently to detect short insertions and deletions in HLA-aligned with increased sensitivity. Both single nucleotide mutations and frameshift alterations passing quality control were annotated by shell_annotate_hla_mutations. Based on this annotation, a mutation in the HLA locus was called if a mutation passed all quality filters and introduced either a missense\/nonsense change or was located at a splice site. In addition, we also identified second-tier mutations in unfiltered MuTect files that were detected in insufficient reads to pass quality control, but the exact same nucleotide change was clearly detected in another biopsy of the same tumor, allowing detection in samples with low purity or sequencing coverage.<\/p>\n<p>LOH in HLA<\/p>\n<p>LOH at the HLA locus was predicted using LOHHLA<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"McGranahan, N. et al. Allele-Specific HLA loss and immune escape in lung cancer evolution. Cell 171, 1259&#x2013;1271 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR33\" id=\"ref-link-section-d101659981e2210\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a> and sequenza<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 49\" title=\"Favero, F. et al. Sequenza: allele-specific copy number and mutation profiles from tumor sequencing data. Ann. Oncol. 26, 64&#x2013;70 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR49\" id=\"ref-link-section-d101659981e2214\" rel=\"nofollow noopener\" target=\"_blank\">49<\/a>. First, we evaluated the allele-specific copy number as predicted by sequenza (v.2.1.2) at the HLA-A, HLA-B and HLA-C loci. Samples with a predicted minor allele copy number of 0 (for example, 2:0, 3:0) were labeled as candidate LOH. Then, we ran LOHHLA with the polysolver-generated haplotype files and matched tumor\u2013normal pairs as input. A type\u2009I allele of a patient was annotated as \u2018allelic imbalance\u2019 (AI) if the associated P\u2009value was lower than 0.01. Alleles with AI were labeled as LOH if the following criteria held: (1) the predicted copy number of the lost allele was below 0.5 with confidence interval (CI) strictly below 0.7; (2) the copy number of the kept allele was above 0.75; (3) the number of mismatched sites between alleles was above 10.<\/p>\n<p>HLA LOH was identified by both methods independently in 18 deep-sequenced samples. HLA LOHs called by only one of the methods were inspected manually. Of LOHHLA-exclusive calls, two were found to be false positives and four were identified as AIs with a minor allele copy number of 1. Of sequenza-exclusive calls, one region showed clear LOH and got classified as HLA LOH; six samples showed similar CN pattern but were missed by LOHHLA due to low purity\u2014as adjacent low-pass sequenced samples showed CN\u2009=\u20091 in the HLA region, we also classified these as HLA LOH. In addition, 25 regions had allelic imbalance detected by LOHHLA that were also confirmed in sequenza to have a CN state of 2:1 or 3:1.<\/p>\n<p>Mutations in APGs<\/p>\n<p>We assembled a list of genes involved in class\u2009I type MHC presentation, by using the KEGG pathway \u2018antigen processing and presentation\u2019 and MHC\u2009I pathway specifically. The following genes were considered: TAP1, TAP2, IRF1, NLRC5, TBK1, PSME3, PSME1, ERAP2, ERAP1, HSPBP1, CALR, B2M, PSME2, PSMA7, CANX, CIITA, TAPBP, CREB1, HLA-A, HLA-B, HLA-C, HSP90AA1, HSP90AB1, HSPA2, HSPA4, HSPA5, HSPA6, HSPA8, IFNG, NFYA, NFYB, NFYC, RFX5, RFXANK and RFXAP. Then, we evaluated the expression of each gene within our cohort and filtered out genes that were not clearly expressed (\u226510 TPM) in at least 5% of samples.<\/p>\n<p>Then, we called evaluated the mutations called in these genes following. Only mutations with at least moderate predicted impact were called.<\/p>\n<p>Neoantigen prediction and proportional burden computation<\/p>\n<p>We predicted neoantigens from somatic mutation calls and patient-specific HLA haplotypes using NeoPredPipe<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 26\" title=\"Schenck, R. O., Lakatos, E., Gatenbee, C., Graham, T. A. &amp; Anderson, A. R. A. NeoPredPipe: high-throughput neoantigen prediction and recognition potential pipeline. BMC Bioinf. 20, 264 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR26\" id=\"ref-link-section-d101659981e2372\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a> for both FF\u2013WGS and FFPE\u2013PS samples. We defined neoantigen burden in a sample as the number of (unique) mutations giving rise to at least one strong-binding (rank &lt;0.5) neoantigen. We focused our analysis on SNV neoantigen burden, unless stated otherwise. We also computed the total protein-changing mutation burden, and used this value to obtain the proportional neoantigen burden for each sample, that is, what percentage of mutations that has the potential to create a neoantigen does actually lead to strong-binder neoantigens.<\/p>\n<p>Defining clonal\/subclonal neoantigens<\/p>\n<p>We assigned clonal\/subclonal categories to all mutations (independent of neoantigen status) based on their presence\/absence in all available deep- or panel-sequenced sample of a given cancer. As the targeted genome region, sequencing strategy and sample types were different, we created separate mutations lists of FF\u2013WGS and FFPE\u2013PS samples. For FF\u2013WGS samples, mutations present in all sequenced cancer biopsies were denoted as clonal and all other mutations (absent in at least one biopsy) as subclonal. For FFPE\u2013PS samples, mutations present in all biopsies were deemed clonal, and mutations absent in at least two biopsies were denoted as subclonal.<\/p>\n<p>Immune dNdS<\/p>\n<p>Immune dNdS was computed using SOPRANO<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Zapata, L. et al. Immune selection determines tumor antigenicity and influences response to checkpoint inhibitors. Nat. Genet. 55, 451&#x2013;460 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR25\" id=\"ref-link-section-d101659981e2392\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>, with somatic mutation files and personalized immunopeptidome files derived specifically to HLA haplotypes. First, the ratio between dNdS inside (ON-target dNdS) and outside the immunopeptidome (OFF-target dNdS) was computed and corrected for 192-trinucleotide context. Then immune dNdS was computed as the ratio of ON-to-OFF dNdS to correct for technical artefacts that could bias dNdS as computed in OFF-target regions. Samples without any ON- or OFF-target synonymous mutations were excluded from the analysis. In total, immune dNdS estimate was available for 61 FF\u2013WGS and 41 FFPE\u2013PS samples.<\/p>\n<p>To compute immune dNdS separately, we first filtered somatic mutation files to only contain mutations annotated as clonal, then repeated the above procedure on these files.<\/p>\n<p>SCAA and SCAA loss analysis<\/p>\n<p>We identified SCAAs by comparing purity-corrected and copy number-corrected ATAC-seq peak calls of cancer regions (per cancer) to a pool of normal glands<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Heide, T. et al. The co-evolution of the genome and epigenome in colorectal cancer. Nature 611, 733&#x2013;743 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR15\" id=\"ref-link-section-d101659981e2407\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>. For immune escape genes, we filtered SCAAs for those located in promoter or enhancer regions associated with a gene from the list detailed in \u2018Mutations in antigen-presenting associated genes.\u2019 Each SCAA was labeled as loss (fold change of cancer compared to normal\u2009&lt;\u2009\u22121) or gain (fold change of cancer compared to normal\u2009&gt;\u20091).<\/p>\n<p>To evaluate the observed number of SCAA losses, we repeated the analysis 200 times with a set of 25\u201340 randomly chosen genes and computed the number of SCAA losses and gains. We derived a P\u2009value as the number of random samples we observed a value more extreme than for APGs.<\/p>\n<p>For neoantigen SCAA loss analysis, we filtered SCAA calls to obtain only losses located in promoter regions (within 1,000\u2009bp of the transcription start site of a gene). Each SCAA loss was annotated by the gene to which they were proximal. As most SCA alterations were found to be clonal, we defined SCAA losses on the cancer level.<\/p>\n<p>We then evaluated for each protein-alteration mutation within a cancer whether the gene it is in had an associated SCAA loss or not. For each cancer, we computed the proportion of mutations (both for neoantigens and for nonantigenic mutations) that were classified as downstream of a SCAA loss. Similarly, for each SCAA loss in a given cancer, we evaluated whether it was upstream of a protein-changing mutation and, if so, within each cancer, we counted the number of SCAA losses upstream of neoantigen\/nonantigenic mutations.<\/p>\n<p>TF analysis<\/p>\n<p>We used the TF binding sites obtained from curated human TF motifs in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Gai, X.-D., Li, C., Song, Y., Lei, Y.-M. &amp; Yang, B.-X. In situ analysis of FOXP3+ regulatory T cells and myeloid dendritic cells in human colorectal cancer tissue and tumor-draining lymph node. Biomed. Rep. 1, 207&#x2013;212 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR21\" id=\"ref-link-section-d101659981e2431\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>, also available as part of the processed Mendeley dataset (\u2018Data availability\u2019).<\/p>\n<p>We examined the list of ATAC-seq peaks that showed statistically significant somatic loss in at least one cancer (27 peaks, representing 20 unique APGs). We then filtered this set to peaks with loss in more than one patient (recurrent SCAA losses), leading to a total of ten genomic regions associated with eight APGs. We identified TFs that are predicted to bind to these regions as those with a binding site in &lt;1,000\u2009bp distance of a given peak, using the distanceToNearest function of the GenomicRanges package in R. We selected for plotting ten TFs that bound more than two of the regions (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>).<\/p>\n<p>Transcriptional editing analysis<\/p>\n<p>For each SNV detected in FF\u2013WGS samples, we quantified the number of reads in the matched RNA-seq data supporting the mutation using bam-readcount (v.1.0.1)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Khanna, A. et al. Bam-readcount - rapid generation of basepair-resolution sequence metrics. J. Open Source Softw. 7, 3722 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR50\" id=\"ref-link-section-d101659981e2450\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a>. For a mutation to be classified as expressed in a sample, we required three or more RNA reads overlapping the position to support the variant base. For a mutation to be classified as not expressed (that is transcriptionally edited), we required more than ten overlapping reads, with zero supporting the variant. We chose the threshold of more than ten so that the probability of misclassifying a mutation present at a true allele frequency of 0.25 or higher was &lt;5%. Mutation\u2013sample pairs that did not qualify for either of these categories were left blank to signify insufficient evidence.<\/p>\n<p>For each cancer, we computed the proportion of mutations that had evidence of transcriptional editing (present in WGS but not expressed in RNA-seq) in at least one sample. Similarly, for each mutation in a cancer, we computed the number of samples that expressed that mutation and the number of samples that (confidently) did not express it. We excluded mutations that had sufficient evidence for expressed\/not status in fewer than two samples. We used the median proportion of biopsies with evidence of editing to derive a single value per cancer per mutation type.<\/p>\n<p>Analysis of H&amp;E slides<\/p>\n<p>Sections were cut from diagnostic FFPE blocks sampled at resection. H&amp;E slides pre- and post-laser capture microdissection (LCM) and post-CyCIF were scanned using the NanoZoomer S210 slide scanner (Hamamatsu). Representative sections were selected for annotations, which were drawn on pre-LCM H&amp;E slides as first choice in most cases. Annotations included: all tumor (on a slide); the tumor\u2013normal interface; any deposits of cancer within nodes; and adjacent normal mucosa (within 5\u2009mm of superficial tumor) and normal mucosa further away (&gt;5\u2009mm away from tumor).<\/p>\n<p>Annotations were made using the NDPViewer software (v.2.9.29). The tumor\u2013normal interface was expanded by 500\u2009\u03bcm on either side to identify the invasive margin. Images with annotations were analyzed using a digital cell classifier, which uses deep learning methodology modeled on a spatially constrained neural network architecture<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Sirinukunwattana, K. et al. in Patch-Based Techniques in Medical Imaging (eds Wu, G. et al.) 154&#x2013;162 (Springer International Publishing, 2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR51\" id=\"ref-link-section-d101659981e2468\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>. The model first detects cells through predicted location of cell nuclei, then classifies them as: normal epithelial, cancer epithelial, fibroblast, lymphocyte, neutrophil, macrophage, endothelial. Absolute counts are calculated for each annotated region. For each annotation, number of cells were normalized to the total number of epithelial cells to obtain per epithelial cell counts.<\/p>\n<p>In addition, we used a further classifier<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Yuan, Y. Modelling the spatial heterogeneity and molecular correlates of lymphocytic infiltration in triple-negative breast cancer. J. R. Soc. Interface 12, 20141153 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR37\" id=\"ref-link-section-d101659981e2475\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a> to class all tumor-associated lymphocytes as infiltrating, adjacent or distant based on proximity to tumor cells.<\/p>\n<p>CyCIF imaging<\/p>\n<p>ROIs that matched the cancer glands\/microbiopsies used for LCM were identified using the H&amp;Es pre-LCM and post-LCM and extended by an additional 1\u2009mm at the edges. All CyCIF sections were 5\u2009\u03bcm thick. The maximum separation between the LCM slides and the CyCIF slide was 10\u2009\u03bcm.<\/p>\n<p>The protocol was based on previous methods from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"Lin, J.-R. et al. Highly multiplexed immunofluorescence imaging of human tissues and tumors using t-CyCIF and conventional optical microscopes. eLife 7, e31657 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#ref-CR24\" id=\"ref-link-section-d101659981e2490\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a>. Full details of fluorescent antibodies are listed in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. Image and statistical analysis of CyCIF images is detailed in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>Statistical analysis<\/p>\n<p>All statistical analyses were carried out in R (v.4.4.2). All comparisons between pairs of samples were carried out using two-sided unpaired Wilcoxon rank sum test, without additional adjustment of P\u2009values, unless stated otherwise. In all figures, visual elements of the boxplots correspond to the following summary statistics: center line, median; box limits, upper and lower quartiles; whiskers, 1.5\u00d7 interquartile range.<\/p>\n<p>Mixed-effect model with patient effect<\/p>\n<p>To account for multiple samples originating from the same patient, we used a mixed-effects model implemented with the R package lme4 (v.1.1.36). We incorporated Patient as a random effect and the tested variable (for example CRA\/CRC) as a fixed effect affecting the intercept into the mixed-effects model. The significance of the fixed effect was evaluated by comparing this \u2018full\u2019 model to one with only random effects, using ANOVA. The P\u2009value of this test (Pr(&gt;Chisq)) is reported as P(fixed effect).<\/p>\n<p>Evaluating neoantigens in expressed genes and transcriptionally\/epigenetically silenced neoantigens<\/p>\n<p>We created 2\u2009\u00d7\u20092 contingency tables by counting mutations that are neoantigens\/nonantigenic and (1) in a gene in the given gene group or not; (2) in a consistently expressed gene or not; (3) in a gene with somatically closed promoter or not or (4) not expressed (missing) in at least one sample or not. Fisher\u2019s exact test was used to compute the OR and CIs on these tables. The above steps were repeated separately on clonal and subclonal mutations alone.<\/p>\n<p>Multivariable regression<\/p>\n<p>Multivariable regression models were constructed using the functions betareg (for proportional burden) and lm (for immune dNdS). Sample type (gland\/bulk for FF\u2013WGS and superficial tumor\/IM\/node for FFPE\u2013PS), tissue type (adenoma\/cancer), immune escape (escape\/weak-escape\/no-escape) and patient were encoded as categorical variables; purity as a continuous variable. For Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>, sample type and patient were encoded as categorical variables and all immune infiltrate values handled as continuous variables (unit: number per epithelial cell). Results were visualized using the plot_summs function from the package jtools, omitting the full list of coefficients assigned to patients (as well as sample type for Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#Fig16\" rel=\"nofollow noopener\" target=\"_blank\">10a<\/a>) to make the visuals easier to interpret. Instead, on Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5a,b<\/a>, the smallest and largest patient-associated coefficients were identified and plotted to highlight the range of coefficients.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02349-1#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Sample collection and sequencing Our FF\u2013WGS samples were comprised of processed data from previous sequencing experiments of our&hellip;\n","protected":false},"author":2,"featured_media":246738,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[5083,5085,3251,847,5082,32937,28560,59,5084,3250,12127,102,5081,56,54,55],"class_list":["post-246737","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-agriculture","tag-animal-genetics-and-genomics","tag-biomedicine","tag-cancer","tag-cancer-research","tag-epigenomics","tag-gastrointestinal-cancer","tag-gb","tag-gene-function","tag-general","tag-genome-informatics","tag-health","tag-human-genetics","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/246737","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=246737"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/246737\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/246738"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=246737"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=246737"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=246737"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}