{"id":19167,"date":"2025-07-24T18:39:31","date_gmt":"2025-07-24T18:39:31","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/19167\/"},"modified":"2025-07-24T18:39:31","modified_gmt":"2025-07-24T18:39:31","slug":"detecting-and-quantifying-clonal-selection-in-somatic-stem-cells","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/19167\/","title":{"rendered":"Detecting and quantifying clonal selection in somatic stem cells"},"content":{"rendered":"<p>Ethical approval<\/p>\n<p>Patient samples were collected with informed consent under the Mechanisms of Age-Related Clonal Haematopoiesis (MARCH) Study. Written informed consent was obtained from all participants in accordance with the Declaration of Helsinki. This study was approved by the Yorkshire &amp; The Humber \u2013 Bradford Leeds Research Ethics Committee (REC Ref17:\/YH\/0382).<\/p>\n<p>Study samples<\/p>\n<p>Participants were recruited from individuals undergoing elective total hip replacement surgery at the Nuffield Orthopaedic Centre, Oxford. Exclusion criteria were history of rheumatoid arthritis or other inflammatory arthritis, history of septic arthritis in the limb undergoing surgery, history of hematological cancer, bisphosphonate use and oral steroid use. Patient characteristics are summarized in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. At the time of surgery, trabecular bone fragments and BM aspirates were obtained from the femoral canal and collected in anticoagulated buffer containing acid\u2013citrate\u2013dextrose, heparin sodium and DNase. Samples of peripheral blood were collected in EDTA vacutainers. Hair follicle samples were collected from participants as a germline control. Peripheral blood and BM mononuclear cells (MNCs) were isolated by Ficoll density gradient centrifugation and viably frozen. Peripheral blood granulocytic cell pellets were frozen for later DNA extraction. Genomic DNA was extracted from BM MNCs, peripheral blood granulocytes and hair follicles using a DNeasy Blood &amp; Tissue Kit (Qiagen).<\/p>\n<p>Cell sorting<\/p>\n<p>Thawing media was prepared with IMDM medium (Gibco) supplemented with 20% fetal bovine serum (FBS) and 110\u2009\u00b5g\u2009ml\u22121 DNase. BM samples were thawed at 37\u2009\u00b0C in a water bath, 1\u2009ml of warm FBS was added and the suspension was then diluted by dropwise addition of 8\u2009ml of thawing media. The suspension was centrifuged at 400\u2009g for 10\u2009min, cells were resuspended in flow cytometry staining medium (IMDM with 10% FBS and 10\u2009\u03bcg\u2009ml\u22121 DNase), filtered through a 35-\u03bcm cell strainer and placed on ice.<\/p>\n<p>Cells were stained with the following antibodies: anti-CD34-PE (1:160, BioLegend, clone 581), anti-CD3-PE\/Cy7 (1:100, BioLegend, clone HIT3a), anti-CD2-PE\/Cy5 (1:160, BioLegend, clone RPA-2.10), anti-CD4-PE\/Cy5 (1:160, BioLegend, clone RPA-T4), anti-CD7-PE\/Cy5 (1:160, BioLegend, clone CD7-6B7), anti-CD8a-PE\/Cy5 (1:320, BioLegend, clone RPA-T8), anti-CD11b-PE\/Cy5 (1:160, BioLegend, clone ICRF44), anti-CD14-PE\/Cy5 (1:160, eBioscience, clone 61D3), anti-CD19-PE\/Cy5 (1:160, BioLegend, clone HIB19), anti-CD20-PE\/Cy5 (1:160, BioLegend, clone 2H7), anti-CD56-PE\/Cy5 (1:80, BioLegend, clone MEM188) and anti-CD235ab-PE\/Cy5 (1:320, BioLegend, clone HIR2). Following antibody incubations, cells were washed with 1\u2009ml of flow cytometry staining buffer, centrifuged at 350\u2009g for 5\u2009min and resuspended in flow cytometry staining buffer containing 1:10,000 Hoechst 33342 live\u2013dead stain.<\/p>\n<p>Cell sorting was performed on a BD FACSAria Fusion or Sony MA900 equipped with a 100-\u00b5m nozzle or sorting chip. Unstained, single-stained and fluorescence-minus-one controls were used to determine background staining and compensation in each channel. Doublets and dead cells were excluded. The following populations were sorted with a mean purity &gt;95%: Lin\u2013CD34+ HSPCs, Lin+CD3+ T cells and Lin+\/\u2013CD34\u2013CD3\u2013 MNCs. Genomic DNA from sorted cell populations was extracted using a QIAamp DNA Micro Kit (Qiagen).<\/p>\n<p>Whole-genome sequencingLibrary preparation and sequencing<\/p>\n<p>Sequencing libraries were prepared using the Illumina DNA Prep Kit (cat. no. 20018704) and Nextera DNA CD Indexes (cat. no. 20018707) according to manufacturer\u2019s guidelines. The Qubit HS DNA Assay Kit (Invitrogen, cat. no. Q32854) and a tape station run with the Agilent HS D5000 Assay Kit (cat. no. 5067-5593) were used for quality control. Thereafter, libraries were diluted to 10\u2009nM, pooled and sequenced on either HiSeq X PE150 or on NovaSeq 6000 PE150.<\/p>\n<p>Low-quality bases on raw sequencing reads were trimmed using Trim Galore v.0.5.0 (<a href=\"https:\/\/www.bioinformatics.babraham.ac.uk\/projects\/trim_galore\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.bioinformatics.babraham.ac.uk\/projects\/trim_galore\/<\/a>) with cutadapt v.2.8 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 60\" title=\"Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J 17, 10&#x2013;12 (2011).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR60\" id=\"ref-link-section-d19200088e2797\" rel=\"nofollow noopener\" target=\"_blank\">60<\/a>) with the following settings: &#8211;quality 30, &#8211;illumina &#8211;length 32 &#8211;trim-n &#8211;clir_R1 2 &#8211;clip_R2 2 &#8211;three_prime_clip_R1 2 &#8211;three_prime_clip_R2 2. Trimmed read pairs were mapped to the human reference genome (build 37 with the GRCh37.75 genome annotation) using bwa mem v.0.7.12 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 61\" title=\"Li, H. &amp; Durbin, R. Fast and accurate short read alignment with Burrows&#x2013;Wheeler transform. Bioinformatics 25, 1754&#x2013;1760 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR61\" id=\"ref-link-section-d19200088e2801\" rel=\"nofollow noopener\" target=\"_blank\">61<\/a>). Mapped reads were coordinate-sorted with samtools sort v.1.5 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 61\" title=\"Li, H. &amp; Durbin, R. Fast and accurate short read alignment with Burrows&#x2013;Wheeler transform. Bioinformatics 25, 1754&#x2013;1760 (2009).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR61\" id=\"ref-link-section-d19200088e2805\" rel=\"nofollow noopener\" target=\"_blank\">61<\/a>), followed by marking duplicate read pairs with gatk MarkDuplicates v.4.0.9.0 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 62\" title=\"Van der Auwera, G.A. &amp; O&#x2019;Connor, B.D. Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O&#x2019;Reilly Media, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR62\" id=\"ref-link-section-d19200088e2809\" rel=\"nofollow noopener\" target=\"_blank\">62<\/a>) and indexing with samtools index.<\/p>\n<p>Detection of SSNVs and insertions\u2013deletions<\/p>\n<p>Somatic SNVs and small insertions\u2013deletions (indels) in BM and blood samples were called with Strelka v.2.9.2 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 63\" title=\"Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat. Methods 15, 591&#x2013;594 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR63\" id=\"ref-link-section-d19200088e2821\" rel=\"nofollow noopener\" target=\"_blank\">63<\/a>) and Mutect2, GATK v.4.2.0.0 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 62\" title=\"Van der Auwera, G.A. &amp; O&#x2019;Connor, B.D. Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O&#x2019;Reilly Media, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR62\" id=\"ref-link-section-d19200088e2825\" rel=\"nofollow noopener\" target=\"_blank\">62<\/a>), using the matched hair follicle samples as the germline control. Variants in repeat regions and simple repeat regions (downloaded from UCSC table browser, setting the assembly to hg19, the track to \u2018RepeatMasker\u2019 or \u2018Simple Repeats\u2019; accession date: 5 December 2018) were filtered using bedtools intersect v.2.24.0 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 64\" title=\"Quinlan, A. R. &amp; Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841&#x2013;842 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR64\" id=\"ref-link-section-d19200088e2829\" rel=\"nofollow noopener\" target=\"_blank\">64<\/a>). Only variants that passed the default filters of both Strelka and Mutect2 were retained. To identify remaining germline variants, variants were looked up in dbSNP (v.150) and the population frequency of reported variants was annotated based on the Genome Aggregation Database (gnomAD v.2.1.1; <a href=\"https:\/\/gnomad.broadinstitute.org\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/gnomad.broadinstitute.org\/<\/a>). Variants were then filtered to exclude potential contamination of germline variants based on the global allele frequency (AF) of gnomAD (retaining variants with AF\u2009&lt;\u20090.001). In a few cases, a CH driver was identified by panel sequencing at low VAF, but not recovered by Strelka and Mutect; here, we re-examined the respective position using bcftools mpileup v.1.10.2 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 65\" title=\"Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR65\" id=\"ref-link-section-d19200088e2840\" rel=\"nofollow noopener\" target=\"_blank\">65<\/a>). All variants were annotated with ANNOVAR (v.May2018; <a href=\"http:\/\/annovar.openbioinformatics.org\/\" rel=\"nofollow noopener\" target=\"_blank\">http:\/\/annovar.openbioinformatics.org\/<\/a>)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 66\" title=\"Wang, K., Li, M. &amp; Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 38, e164&#x2013;e164 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR66\" id=\"ref-link-section-d19200088e2852\" rel=\"nofollow noopener\" target=\"_blank\">66<\/a> according to the human genome build v.19. For SSNVs, VAFs were recalculated directly from the BAM files using alleleCounter v.4.0.2 (<a href=\"https:\/\/github.com\/cancerit\/alleleCount\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/cancerit\/alleleCount<\/a>; with default base and mapping quality thresholds) and MaC (<a href=\"https:\/\/github.com\/nansari-pour\/MaC\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/nansari-pour\/MaC<\/a>).<\/p>\n<p>Detection of copy number variants<\/p>\n<p>Genome-wide subclonal CNAs were identified with Battenberg (v.2.2.10; <a href=\"https:\/\/github.com\/Wedge-lab\/battenberg\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/Wedge-lab\/battenberg<\/a>), which has been described in detail previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Nik-Zainal, S. et al. The life history of 21 breast cancers. Cell 149, 994&#x2013;1007 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR67\" id=\"ref-link-section-d19200088e2885\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a>.<\/p>\n<p>Detection of structural variants<\/p>\n<p>SVs from whole-genome data were identified using Manta v.1.6.0 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 68\" title=\"Chen, X. et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics 32, 1220&#x2013;1222 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR68\" id=\"ref-link-section-d19200088e2897\" rel=\"nofollow noopener\" target=\"_blank\">68<\/a>) with default tumor\u2013normal pair settings comparing samples with the germline controls. We kept variants that passed Manta\u2019s default filter, had a minimal SOMATICSCORE of 30, were not classified as IMPRECISE, had at least three variants and a VAF of 5% in the blood sample, and at most two variant reads and VAF\u2009&lt;\u20095% in the control sample. SVs previously described in somatic tissues<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 57\" title=\"Vattathil, S. &amp; Scheet, P. Extensive hidden genomic mosaicism revealed in normal tissue. Am. J. Hum. Genet. 98, 571&#x2013;578 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR57\" id=\"ref-link-section-d19200088e2901\" rel=\"nofollow noopener\" target=\"_blank\">57<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Laurie, C. C. et al. Detectable clonal mosaicism from birth to old age and its relationship to cancer. Nat. Genet. 44, 642&#x2013;650 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR58\" id=\"ref-link-section-d19200088e2904\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a> and associated with hematological cancer<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Loh, P.-R. et al. Insights into clonal haematopoiesis from 8,342 mosaic chromosomal alterations. Nature 559, 350&#x2013;355 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR21\" id=\"ref-link-section-d19200088e2908\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 58\" title=\"Laurie, C. C. et al. Detectable clonal mosaicism from birth to old age and its relationship to cancer. Nat. Genet. 44, 642&#x2013;650 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR58\" id=\"ref-link-section-d19200088e2911\" rel=\"nofollow noopener\" target=\"_blank\">58<\/a> were classified as putative drivers (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>).<\/p>\n<p>Driver analysis<\/p>\n<p>To comprehensively search for candidate driver mutations, we concatenated published lists of putative driver genes in CH and leukemia from Intogen<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 69\" title=\"Gonzalez-Perez, A. et al. IntOGen-mutations identifies cancer drivers across tumor types. Nat. Methods 10, 1081&#x2013;1082 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR69\" id=\"ref-link-section-d19200088e2927\" rel=\"nofollow noopener\" target=\"_blank\">69<\/a> (subsetting genes described in \u2018ALL\u2019, \u2018AML\u2019, \u2018CLL\u2019, \u2018CML\u2019 or \u2018MM\u2019 from release 2020.02.01, and the CH gene list from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 70\" title=\"Pich, O., Reyes-Salazar, I., Gonzalez-Perez, A. &amp; Lopez-Bigas, N. Discovering the drivers of clonal hematopoiesis. Nat. Commun. 13, 4267 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR70\" id=\"ref-link-section-d19200088e2931\" rel=\"nofollow noopener\" target=\"_blank\">70<\/a>), Cosmic<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 71\" title=\"Sondka, Z. et al. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat. Rev. Cancer 18, 696&#x2013;705 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR71\" id=\"ref-link-section-d19200088e2935\" rel=\"nofollow noopener\" target=\"_blank\">71<\/a> v.94 (subsetting on genes with annotation in leukemia), ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Mitchell, E. et al. Clonal dynamics of haematopoiesis across the human lifespan. Nature 606, 343&#x2013;350 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR13\" id=\"ref-link-section-d19200088e2939\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>, and a curated list of CH-associated mutations compiled from five large studies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Jaiswal, S. et al. Age-related clonal hematopoiesis associated with adverse outcomes. N. Engl. J. Med. 371, 2488&#x2013;2498 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR19\" id=\"ref-link-section-d19200088e2943\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Genovese, G. et al. Clonal hematopoiesis and blood-cancer risk inferred from blood DNA sequence. N. Engl. J. Med. 371, 2477&#x2013;2487 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR20\" id=\"ref-link-section-d19200088e2946\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Jakobsen, N. A. et al. Selective advantage of mutant stem cells in human clonal hematopoiesis is associated with attenuated response to inflammation and aging. Cell Stem Cell 31, 1127&#x2013;1144 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR27\" id=\"ref-link-section-d19200088e2949\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Abelson, S. et al. Prediction of acute myeloid leukaemia risk in healthy individuals. Nature 559, 400&#x2013;404 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR30\" id=\"ref-link-section-d19200088e2952\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Acuna-Hidalgo, R. et al. Ultra-sensitive sequencing identifies high prevalence of clonal hematopoiesis-associated mutations throughout adult life. Am. J. Hum. Genet. 101, 50&#x2013;64 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR72\" id=\"ref-link-section-d19200088e2955\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 73\" title=\"Desai, P. et al. Somatic mutations precede acute myeloid leukemia years before diagnosis. Nat. Med. 24, 1015&#x2013;1023 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR73\" id=\"ref-link-section-d19200088e2958\" rel=\"nofollow noopener\" target=\"_blank\">73<\/a>.<\/p>\n<p>Mutations identified by Strelka and targeting any of these genes were looked up in ClinVar<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 74\" title=\"Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062&#x2013;D1067 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR74\" id=\"ref-link-section-d19200088e2965\" rel=\"nofollow noopener\" target=\"_blank\">74<\/a> (v.20221231) and annotated accordingly. Using the R package drawProteins<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 75\" title=\"Brennan, P. drawProteins: a Bioconductor\/R package for reproducible and programmatic generation of protein schematics. F1000Research 7, 1105 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR75\" id=\"ref-link-section-d19200088e2969\" rel=\"nofollow noopener\" target=\"_blank\">75<\/a> v.1.18.0 we collected information on the protein domain targeted by each variant. Moreover, we manually annotated variant information (involvement in disease, functional evidence for mutations at this site, SNP identifier (SNPID), association with genetic disorders and whether the variant targets a functional domain) using the manually curated variant information \u2018homo_sapiens_variation.txt\u2019 from Uniprot<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 76\" title=\"Consortium, U. UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res. 47, D506&#x2013;D515 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR76\" id=\"ref-link-section-d19200088e2973\" rel=\"nofollow noopener\" target=\"_blank\">76<\/a> (downloaded 4 April 2023) and variant information provided on <a href=\"http:\/\/www.uniprot.org\" rel=\"nofollow noopener\" target=\"_blank\">www.uniprot.org<\/a>. Thereafter, we computed SIFT<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 77\" title=\"Ng, P. C. &amp; Henikoff, S. SIFT: predicting amino acid changes that affect protein function. Nucleic Acids Res. 31, 3812&#x2013;3814 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR77\" id=\"ref-link-section-d19200088e2984\" rel=\"nofollow noopener\" target=\"_blank\">77<\/a> and Revel<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"Ioannidis, N. M. et al. REVEL: an ensemble method for predicting the pathogenicity of rare missense variants. Am. J. Hum. Genet. 99, 877&#x2013;885 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR78\" id=\"ref-link-section-d19200088e2989\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a> prediction scores for each substitution. We kept variants causing a \u2018stopgain\u2019, \u2018stoploss\u2019, \u2018frameshift_insertion\u2019 or \u2018frameshift_deletion\u2019, or with \u2019Conflicting interpretations of pathogenicity\u2019, \u2018Likely pathogenic\u2019, \u2018Pathogenic\/Likely pathogenic\u2019 or of \u2018uncertain significance\u2019 according to ClinVar, or annotated to a related disease (based on uniprot.org) or annotated as pathogenic (uniprot.org) or with a REVEL score \u22650.75 or targeting either ASXL1, or DNMT3A or TET2.<\/p>\n<p>This yielded a set of variant positions, which we merged with curated positions in known CH drivers<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Jakobsen, N. A. et al. Selective advantage of mutant stem cells in human clonal hematopoiesis is associated with attenuated response to inflammation and aging. Cell Stem Cell 31, 1127&#x2013;1144 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR27\" id=\"ref-link-section-d19200088e3005\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>. Based on this set, we re-called variants in all blood and control samples with bcftools mpileup<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 65\" title=\"Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR65\" id=\"ref-link-section-d19200088e3009\" rel=\"nofollow noopener\" target=\"_blank\">65<\/a> v.1.10.2, using the option -p, and bcftools call, using the option -mA. Variant calls were performed in batches and merged using bcftools merge. In the final set, we kept variants targeting any of ASXL1, DNMT3A or TET2 and variants that were found on at least three reads in the sample and on fewer than five reads in the controls.<\/p>\n<p>Analysis of single-cell WGS data<\/p>\n<p>We tested our population genetics model on published single-cell WGS data from refs. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Lee-Six, H. et al. Population dynamics of normal human blood inferred from somatic mutations. Nature 561, 473&#x2013;478 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR2\" id=\"ref-link-section-d19200088e3032\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Fabre, M. A. et al. The longitudinal dynamics and natural history of clonal haematopoiesis. Nature 606, 335&#x2013;342 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR12\" id=\"ref-link-section-d19200088e3035\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Mitchell, E. et al. Clonal dynamics of haematopoiesis across the human lifespan. Nature 606, 343&#x2013;350 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR13\" id=\"ref-link-section-d19200088e3038\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>. To this end, we computed pseudo-bulk VAFs from the single-cell phylogenies as<\/p>\n<p>$${\\text{VAF}}=\\frac{{n}_{\\text{variant cells}}}{2{n}_{\\text{cells}}}$$<\/p>\n<p>where nvariant cells is the number of cells harboring a variant and ncells is the total number of sequenced cells; the factor 2 accounts for diploidy.<\/p>\n<p>Parameter estimation<\/p>\n<p>We fit our population genetics model to the cumulative VAF distribution truncated at 0.01, as detailed below and using the prior probabilities outlined in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>.<\/p>\n<p>Reanalysis of SSNV and indels<\/p>\n<p>To assess differences in variant calling pipelines, we re-called SSNVs and indels in the data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Lee-Six, H. et al. Population dynamics of normal human blood inferred from somatic mutations. Nature 561, 473&#x2013;478 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR2\" id=\"ref-link-section-d19200088e3136\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. To this end, we intersected the results between Mutect2 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 62\" title=\"Van der Auwera, G.A. &amp; O&#x2019;Connor, B.D. Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O&#x2019;Reilly Media, 2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR62\" id=\"ref-link-section-d19200088e3140\" rel=\"nofollow noopener\" target=\"_blank\">62<\/a>) and Strelka2 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 63\" title=\"Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat. Methods 15, 591&#x2013;594 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR63\" id=\"ref-link-section-d19200088e3144\" rel=\"nofollow noopener\" target=\"_blank\">63<\/a>), filtered remaining germline variants using gnomAD (retaining variants with AF &lt;0.001) and recalculated VAFs directly from the BAM files using alleleCounter v.4.0.2 (<a href=\"https:\/\/github.com\/cancerit\/alleleCount\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/cancerit\/alleleCount<\/a>; with default base and mapping quality thresholds) and MaC (<a href=\"https:\/\/github.com\/nansari-pour\/MaC\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/nansari-pour\/MaC<\/a>). We filtered the remaining variants as stated in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Lee-Six, H. et al. Population dynamics of normal human blood inferred from somatic mutations. Nature 561, 473&#x2013;478 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR2\" id=\"ref-link-section-d19200088e3163\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>. In brief, we removed variants occurring in &gt;120 of the 140 colonies, variants that fell within 10\u2009bp of each other and variants with a coverage &lt;6 on autosomes or &lt;3 on sex chromosomes in more than five samples. Moreover, we retained only variants with a mean VAF\u2009&gt;\u20090.3 across all samples with at least one mutant read. Finally, we excluded sites at which &gt;10% of the samples with at least one mutant read had a VAF\u2009&lt;\u20090.1.<\/p>\n<p>Phylogenetic reconstruction<\/p>\n<p>We reconstructed single-cell phylogenies based on the re-called SSNVs and indels from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 2\" title=\"Lee-Six, H. et al. Population dynamics of normal human blood inferred from somatic mutations. Nature 561, 473&#x2013;478 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR2\" id=\"ref-link-section-d19200088e3175\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> by converting the mutation table into a fasta file, learning the phylogenetic tree using MPBoot v.1.1.0 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"Hoang, D. T. et al. MPBoot: fast phylogenetic maximum parsimony tree inference and bootstrap approximation. BMC Evol. Biol. 18, 11 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR79\" id=\"ref-link-section-d19200088e3179\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>) and plotting the tree using custom scripts from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Mitchell, E. et al. Clonal dynamics of haematopoiesis across the human lifespan. Nature 606, 343&#x2013;350 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR13\" id=\"ref-link-section-d19200088e3183\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a> (<a href=\"https:\/\/github.com\/emily-mitchell\/normal_haematopoiesis\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/emily-mitchell\/normal_haematopoiesis\/<\/a>).<\/p>\n<p>Analysis of brain samples<\/p>\n<p>SSNVs of 457 human brain samples were downloaded from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Bae, T. et al. Analysis of somatic mutations in 131 human brains reveals aging-associated hypermutability. Science 377, 511&#x2013;517 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR17\" id=\"ref-link-section-d19200088e3203\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>. Of these, we analyzed 177 samples with an average coverage &gt;100\u00d7, as well as 8 samples from individual NC7 (NC7-CX-ASTMIG, NC7-CX-INT, NC7-CX-OLI, NC7-CX-PYR, NC7-STR-ASTMIG, NC7-STR-INT, NC7-STR-MSN and NC7-STR-OLI) that contained multiple known driver mutations but had average coverage between 32\u00d7 and 40\u00d7 (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). Both tier1 and tier2 variant calls were used for analysis.<\/p>\n<p>Population genetics modelTheory<\/p>\n<p>We modeled the evolution of VAFs mechanistically, accounting for accumulation, drift and selection of somatic variants in a homeostatic tissue. The model is parametrized with the rate at which HSCs divide during adulthood, \u03bb, the number of SSNVs acquired between two cell divisions, \u03bc, the number of HSCs during adulthood, Nss, as well as the time of origin of the selected clone, ts, and its selective advantage expressed by r. We bundled the model functions in an R package, SCIFER, available on <a href=\"https:\/\/github.com\/VerenaK90\/SCIFER\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/VerenaK90\/SCIFER<\/a>.<\/p>\n<p>                    Modeling the site frequency spectrum of somatic variants generated by neutral evolution<\/p>\n<p>The time-dependent site frequency spectrum Si(t) gives the number of variants with clone size i, where i ranges from some minimally observable clone size (owing to WGS sequencing depth) up to the total number of stem cells N(t). We derive analytical expressions for the site frequency spectra (SFS) resulting from neutral evolution and clonal selection and compare these with measured VAF histograms.<\/p>\n<p>To begin with neutral evolution, we develop a stochastic model for accumulation and drift of neutral somatic variants during developmental expansion and subsequent homeostasis of the stem cell pool. The model assumes that stem cells proliferate via symmetric self-renewing divisions, with rate \u03bb, and are lost by differentiation and cell death, with rate \u03b4. Between two subsequent stem cell divisions, on average \u03bc new variants are introduced. These variants are inherited to daughter cells and, depending on the dynamics of the corresponding stem cell clone, may either go extinct or drift to variable frequencies. The SFS generated by these dynamics is:<\/p>\n<p>$${S}_{i}\\left(t\\right)=\\mathop{\\int}\\limits_{0}^{t}\\lambda \\mu N\\left({t}^{{\\prime} }\\right){P}_{1,i}(t,t{\\prime} ){\\rm{d}}{t}^{{\\prime} }$$<\/p>\n<p>describing the generation of new variants at time t\u2032, in a population of size N(t\u2032), and their drift to clone size i up to the time of measurement t with probability P1,i(t,t\u2032). To compare Si(t) with measured VAF histograms, we transform from clone size to VAF:<\/p>\n<p>$${\\text{VAF}}=\\frac{i}{2N}$$<\/p>\n<p>To describe developmental expansion of the stem cell pool followed by homeostasis, we concatenated linear birth\u2013death processes. During development, the division rate will exceed the loss rate, \u03bbexp\u2009&gt;\u2009\u03b4exp, hence defining a supercritical birth\u2013death process. At time t1, the system reaches its steady-state with a constant number of active stem cells, Nss. The cellular dynamics are now appropriately described by a critical birth\u2013death process with steady-state rate \u03bbss\u2009=\u2009\u03b4ss.<\/p>\n<p>To compute the SFS during developmental expansion, we consider the probability that a cell acquiring a new variant will expand to a clone of size a in time t (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Bailey, N. The Elements of Stochastic Processes (Wiley, 1964).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR80\" id=\"ref-link-section-d19200088e3538\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>):<\/p>\n<p>$${P}_{\\exp ,1,a}\\left(t\\right)=\\left\\{\\begin{array}{c}x(t),\\text{if}\\;{a}=0\\\\ \\left(1-x\\left(t\\right)\\right)\\left(1-y\\left(t\\right)\\right){y\\left(t\\right)}^{a-1},\\text{if}\\;a\\ge 1\\end{array}\\right.$$<\/p>\n<p>\n                    (1)\n                <\/p>\n<p>with<\/p>\n<p>$${{x}}\\left({{t}}\\right)=\\displaystyle\\frac{{\\delta }_{\\exp }{{\\rm{e}}}^{({\\lambda }_{\\exp }-{\\delta }_{\\exp }){{t}}}-{\\delta }_{\\exp }}{{\\lambda }_{\\exp }{{\\rm{e}}}^{({\\lambda }_{\\exp }-{\\delta }_{\\exp }){{t}}}-{\\delta }_{\\exp }},{{y}}\\left({{t}}\\right)=\\displaystyle\\frac{{\\lambda }_{\\exp }{{{e}}}^{({\\lambda }_{\\exp }-{\\delta }_{\\exp }){{t}}}-{\\lambda }_{\\exp }}{{\\lambda }_{\\exp }{{{e}}}^{({\\lambda }_{\\exp }-{\\delta }_{\\exp }){{t}}}-{\\delta }_{\\exp }}$$<\/p>\n<p>\n                    (2)\n                <\/p>\n<p>The measured VAF histograms report on the number of variants with a given frequency. To calculate variant number in the model, we note that between time t\u2032 and t\u2032\u2009+\u2009dt\u2032 on average \u03bc\u03bbexpN(t\u2032)dt\u2032 variants are generated. Hence,<\/p>\n<p>$${\\mu {\\lambda }_{\\exp }N\\left({t}^{{\\prime} }\\right)\\text{d}t{\\prime} \\times P}_{\\exp ,1,a}\\left(t-{t}^{{\\prime} }\\right)$$<\/p>\n<p>\n                    (3)\n                <\/p>\n<p>variants introduced at time t\u2032 each occur in a cells at time t.<\/p>\n<p>During tissue homeostasis, when stem cell division and loss will occur both at steady-state rate \u03bbss, drift is described by the critical birth\u2013death process. A variant in a cells at t1 (when the homeostatic stem cell number is reached) will drift to occur in b cells within time t with probability<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Bailey, N. The Elements of Stochastic Processes (Wiley, 1964).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR80\" id=\"ref-link-section-d19200088e4280\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a><\/p>\n<p>$${P}_{\\text{ss},a,b}(t)=\\left\\{\\begin{array}{c}{p(t)}^{a},{\\text{if}}\\,b=0\\\\ \\mathop{\\sum }\\limits_{j=0}^{{{b}}}\\frac{j}{b}{{a}\\choose{j}}{p\\left(t\\right)}^{a-j}{\\left(1-p(t)\\right)\\;}^{j}{{b}\\choose{j}}{p(t)}^{b-j}{\\left(1-p(t)\\right)}{j\\atop},{\\text{otherwise}}\\end{array}\\right.$$<\/p>\n<p>\n                    (4)\n                <\/p>\n<p>with<\/p>\n<p>$$p\\left(t\\right)=\\frac{{\\lambda }_{\\text{ss}}t}{1+{\\lambda }_{\\text{ss}}t}$$<\/p>\n<p>\n                    (5)\n                <\/p>\n<p>Thus, at t\u2009\u2265\u2009t1, the number of variants occurring exactly in b cells is<\/p>\n<p>$$\\mathop{\\sum }\\limits_{a=1}^{{N}_{\\text{ss}}}{\\mu {\\lambda }_{\\exp }N({t}^{{\\prime} })P}_{\\exp ,1,a}\\left({t}_{1}-{t}^{{\\prime} }\\right){P}_{\\text{ss},a,b}\\left(t-{t}_{1}\\right)\\text{d}{t}^{{\\prime} }$$<\/p>\n<p>\n                    (6)\n                <\/p>\n<p>for variants generated during developmental expansion (t\u2032\u2009&lt;\u2009t1).<\/p>\n<p>Finally, we consider variants acquired during homeostasis, which evolve according to the critical birth\u2013death process entirely. Hence, the number of such variants occurring exactly in b cells is<\/p>\n<p>$$\\mu {\\lambda }_{\\text{ss}}{N}_{\\text{ss}}\\times {P}_{\\text{ss},1,b}\\left(t-t{\\prime} \\right)\\text{d}{t}^{{\\prime} }$$<\/p>\n<p>\n                    (7)\n                <\/p>\n<p>where t\u2032\u2009\u2265\u2009t1. Combining the contribution of both phases, we arrive at the SFS of neutral variants in a homeostatic tissue without selection:<\/p>\n<p>$$\\begin{array}{l}{S}_{i}\\left(t\\right)=\\underbrace{{\\displaystyle\\int }_{\\!0}^{{t}_{1}}\\mathop{\\sum }\\limits_{a=1}^{{N}_{\\text{ss}}}{\\mu {\\lambda }_{\\exp }N\\left({t}^{{\\prime} }\\right)P}_{\\exp ,1,a}\\left({t}_{1}-{t}^{{\\prime} }\\right){P}_{\\text{ss},a,i}\\left(t-{t}_{1}\\right)\\text{d}{t}^{{\\prime} }}_{{\\rm{variants}}\\; {\\rm{generated}}\\; {\\rm{in}}\\; {\\rm{development}}}\\\\\\quad\\quad+\\underbrace{{\\displaystyle\\int }_{\\!{t}_{1}}^{t}\\mu {\\lambda }_{\\text{ss}}{N}_{\\text{ss}}{P}_{\\text{ss},1,i}\\left(t-{t}^{{\\prime} }\\right)\\text{d}{t}^{{\\prime} }}_{{\\rm{variants}}\\; {\\rm{generated}}\\; {\\rm{in}}\\; {\\rm{homeostasis}}}\\end{array}$$<\/p>\n<p>\n                    (8)\n                <\/p>\n<p>Equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ8\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>) generalizes a result by Ohtsuki and Innan<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 81\" title=\"Ohtsuki, H. &amp; Innan, H. Forward and backward evolutionary processes and allele frequency spectrum in a cancer cell population. Theor. Popul. Biol. 117, 43&#x2013;50 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR81\" id=\"ref-link-section-d19200088e5528\" rel=\"nofollow noopener\" target=\"_blank\">81<\/a> for expanding tissues.<\/p>\n<p>                    Modeling the site frequency spectrum of somatic variants under selection<\/p>\n<p>Modeling a single selected clone in a homeostatic tissue. Clonal selection will modify the SFS. Consider that a positively selected mutation is acquired at time ts and reduces the rate of stem cell loss by the factor r\u2009&lt;\u20091 (the alternative for imparting selective advantage, an increase in the rate of stem cell division, yields very similar results). The cell number of mutated stem cells, n2(t), expands at the expense of the number of normal stem cells, n1(t)\u2009=\u2009Nss\u2009\u2212\u2009n2(t), according to the competition model<\/p>\n<p>$$\\begin{array}{c}\\frac{{\\rm{d}}{n}_{1}}{{{\\rm{d}}t}}={\\lambda }_{\\text{ss}}{n}_{1}\\left(t\\right)\\left(1-{\\rho }_{{n}_{1},{n}_{1}}\\frac{{n}_{1}(t)}{{N}_{\\text{ss}}}-{\\rho }_{{n}_{1},{n}_{2}}\\frac{{n}_{2}\\left(t\\right)}{{N}_{\\text{ss}}}\\right)=g({n}_{1},{n}_{2}),\\\\ \\frac{{\\rm{d}}{n}_{2}}{{{\\rm{d}}t}}={\\lambda }_{\\text{ss}}{n}_{2}\\left(t\\right)\\left(1-{\\rho }_{{n}_{2},{n}_{2}}\\frac{{n}_{2}(t)}{{N}_{\\text{ss}}}-{\\rho }_{{n}_{2},{n}_{1}}\\frac{{n}_{1}\\left(t\\right)}{{N}_{\\text{ss}}}\\right)=h({n}_{1},{n}_{2})\\end{array}$$<\/p>\n<p>\n                    (9)\n                <\/p>\n<p>The \u03c1 parameters denote phenomenological competition coefficients between and in the mutant clone and the normal stem cells, maintaining homeostasis. We have no further growth if either normal or mutated stem cells fill the entire compartment, g(Nss, 0)\u2009=\u20090\u2009=\u2009h(0, Nss) and g(0, Nss)\u2009=\u20090\u2009=\u2009h(Nss, 0), implying that \\({\\rho }_{{n}_{1},{n}_{1}}=1={\\rho }_{{n}_{2},{n}_{2}}\\). Moreover, \\({\\rho }_{{n}_{2},{n}_{1}}=r\\) and hence<\/p>\n<p>$$\\frac{{\\rm{d}}{n}_{1}}{{{\\rm{d}}t}}={\\lambda }_{\\text{ss}}{n}_{1}\\left(t\\right)\\left(1-\\frac{{n}_{1}\\left(t\\right)}{{N}_{\\text{ss}}}-\\left(2-r\\right)\\frac{{n}_{2}\\left(t\\right)}{{N}_{\\text{ss}}}\\right)$$<\/p>\n<p>\n                    (10)\n                <\/p>\n<p>$$\\frac{{\\rm{d}}{n}_{2}}{{{\\rm{d}}t}}={\\lambda }_{\\text{ss}}{n}_{2}\\left(t\\right)\\left(1-\\frac{{n}_{2}(t)}{{N}_{\\text{ss}}}-r\\frac{{n}_{1}\\left(t\\right)}{{N}_{\\text{ss}}}\\right)$$<\/p>\n<p>\n                    (11)\n                <\/p>\n<p>The number of mutant stem cells n2(t) will remain much smaller than the number of normal stem cells for extended periods, and hence we approximate equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ11\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>) by<\/p>\n<p>$$\\frac{{\\rm{d}}{n}_{2}}{{{\\rm{d}}t}}={\\lambda }_{\\text{ss}}{n}_{2}\\left(t\\right)\\left(1-r\\right)$$<\/p>\n<p>\n                    (12)\n                <\/p>\n<p>Equations (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ10\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>) and (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ12\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>) will be used when computing the SFS.<\/p>\n<p>With clonal selection, the SFS has three principal contributions for somatic variants originating: (1) before the positively selected mutation occurred, Si,1; (2) after this mutation occurred and happening in the mutant clone, Si,2; and (3) after the mutation occurred but happening in normal stem cells, Si,3:<\/p>\n<p>$${S}_{i}\\left(t\\right)={S}_{i,1}\\left(t\\right)+{S}_{i,2}\\left(t\\right)+{S}_{i,3}\\left(t\\right)$$<\/p>\n<p>\n                    (13)\n                <\/p>\n<p>We specify each contribution in turn. The SFS in the mutant clone, Si,2, is<\/p>\n<p>$${S}_{i,2}\\left(t\\right)={\\int }_{\\!{t}_{\\text{s}}}^{t}{\\mu {\\lambda }_{\\text{ss}}{n}_{2}({t}^{{\\prime} })P}_{\\exp ,1,i}\\left(t-{t}^{{\\prime} }|{\\lambda }_{\\text{ss}},r{\\lambda }_{\\text{ss}}\\right)\\text{d}{t}^{{\\prime} }$$<\/p>\n<p>\n                    (14)\n                <\/p>\n<p>where Pexp,1,i is given by equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>).<\/p>\n<p>The SFS in normal stem cells, Si,1 and Si,3, are shaped by the decline in the number of normal stem cells after tS (equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ10\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>)). We approximate this process with a subcritical birth\u2013death process with division rate \u03bbss and effective loss rate \u03b4eff(t). The loss rate is chosen such that the expectation of the decline of normal stem cells in the linear birth\u2013death process equals the number of normal stem cells lost by the competition dynamics, D. Equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ10\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>) implies that<\/p>\n<p>$$D={\\int }_{\\!{t}_{\\text{s}}}^{t}{\\lambda }_{\\text{ss}}\\left(\\frac{{n}_{1}\\left(t\\right)}{{N}_{\\text{ss}}}+\\left(2-r\\right)\\frac{{n}_{2}\\left(t\\right)}{{N}_{\\text{ss}}}\\right){n}_{1}\\left(t\\right)\\text{d}t$$<\/p>\n<p>\n                    (15)\n                <\/p>\n<p>Therefore, the effective loss rate \u03b4eff(t) is defined by<\/p>\n<p>$${\\int }_{\\!{t}_{\\text{s}}}^{t}{\\delta }_{\\text{eff}}{N}_{\\text{ss}}\\exp \\left(\\left({\\lambda }_{\\text{ss}}-{\\delta }_{\\text{eff}}\\right)(t-{t}_{\\text{s}})\\right)\\text{d}t=D$$<\/p>\n<p>\n                    (16)\n                <\/p>\n<p>The SFS of variants acquired in normal stem cells after ts, Si,3, is the superposition of (Nss\u2009\u2212\u20091) independent linear subcritical birth\u2013death processes and is given by<\/p>\n<p>$${S}_{i,3}\\left(t\\right)=(N_{\\text{ss}}-1){\\int }_{\\!{t}_{\\text{s}}}^{t}{\\mu {\\lambda }_{\\text{ss}}{\\rm{e}}^{\\left({\\lambda }_{\\text{ss}}-{\\delta }_{\\text{eff}}\\right)(t^{{\\prime} }-{t}_{\\text{s}})}P}_{\\exp ,1,i}\\left(t-{t}^{{\\prime} }|{\\lambda }_{\\text{ss}},{\\delta }_{\\text{eff}}\\right)\\text{d}{t}^{{\\prime} }$$<\/p>\n<p>\n                    (17)\n                <\/p>\n<p>Finally, we compute the SFS of variants acquired before ts, Si,1. These variants may be inherited to the selected clone, in which case they will be present in all selected cells and, additionally, in some of the normal cells. Alternatively, they may be present in normal cells only. To distinguish the two cases, we consider a variant that was acquired before ts and is present in k cells at ts. Assuming that the driver mutation is acquired in a random cell, the probability of this variant being inherited to the selected clone is \\(\\frac{k}{{N}_{\\text{ss}}}\\), while the probability of this variant being exclusively present in normal cells is \\(\\frac{1-k}{{N}_{\\text{ss}}}\\). In the former case, all n2(t) selected cells will harbor the variant at the time of measurement, t; in addition, the k\u2009\u2212\u20091 normal stem cells harboring the variant at time tS may reach a clone size between 0 and Nss\u2009\u2212\u2009n2(t), according to the drift dynamics of normal stem cells. In the alternative case, the variant does not end up in the selected clone and hence solely drifts in the normal cells. Taken together, this yields<\/p>\n<p>$${S}_{i,1}\\left(t\\right)=\\left\\{\\begin{array}{c}\\mathop{\\sum }\\limits_{k=1}^{{N}_{\\text{ss}}}{S}_{k}\\left({t}_{\\text{s}}\\right)\\left[\\overbrace{\\displaystyle\\frac{k}{{N}_{\\text{ss}}}{P}_{\\exp ,k-1,i-{n}_{2}\\left(t\\right)}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}},{\\delta }_{\\text{eff}}\\right)}^{\\rm{variants}\\; \\rm{drift}\\; \\rm{in}\\; \\rm{normal}\\; \\rm{cells}\\;{\\rm{and}}\\;\\rm{present}\\; \\rm{in}\\; \\rm{all}\\; \\rm{selected}\\; \\rm{cells}}\\right.\\\\\\qquad\\qquad\\;\\left.+\\overbrace{\\left(1-\\displaystyle\\frac{k}{{N}_{\\text{ss}}}\\right){P}_{\\exp ,k,i}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}},{\\delta }_{\\text{eff}}\\right)}^{\\rm{variants}\\; \\rm{not}\\; \\rm{present}\\; \\rm{in}\\; \\rm{selected}\\; \\rm{cells}}\\right],\\,i\\ge {n}_{2}\\left(t\\right),\\\\\\qquad\\qquad\\qquad\\; \\mathop{\\sum }\\limits_{k=1}^{{N}_{\\text{ss}}}\\left(1-\\displaystyle\\frac{k}{{N}_{\\text{ss}}}\\right){S}_{k}\\left({t}_{\\text{s}}\\right){P}_{\\exp ,k,i}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}},{\\delta }_{\\text{eff}}\\right),\\,i &lt; {n}_{2}(t),\\end{array}\\right.$$<\/p>\n<p>\n                    (18)\n                <\/p>\n<p>where Sk(ts) is the SFS at ts and Pexp,a,b is the clone size distribution generated by a subcritical birth\u2013death process initiated by a cells<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Bailey, N. The Elements of Stochastic Processes (Wiley, 1964).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR80\" id=\"ref-link-section-d19200088e9084\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a>:<\/p>\n<p>$${P}_{\\exp ,a,b}(t)=\\left\\{\\begin{array}{c}{x\\left(t\\right)}^{a},\\text{if}\\,b=0,\\\\ {\\sum }_{j=0}^{\\min \\left(a,b\\right)}{{\\;a\\;}\\choose{\\;j\\;}}{{a+b-j-1}\\choose{a-1}}{x\\left(t\\right)}^{a-j}{y\\left(t\\right)}^{b-j}{\\left(1-x\\left(t\\right)-y\\left(t\\right)\\right)}^{\\;j},\\text{if}\\,b\\ge 1\\end{array}\\right.$$<\/p>\n<p>\n                    (19)\n                <\/p>\n<p>Thus, we have specified the contributions to the SFS of a stem cell population containing a selected clone, Si(t)\u2009=\u2009Si,1(t)\u2009+\u2009Si,2(t)\u2009+\u2009Si,3(t).<\/p>\n<p>Modeling a single selected clone without size compensation. We here develop a version of our model in which the selected clone expands unrestrictedly. We assume a constant number Nss of normal stem cells, whereas the selected clone exponentially expands. Hence, the population dynamics of normal (n1) and mutant (n2) stem cells now read<\/p>\n<p>$$\\begin{array}{c}{n}_{1}\\left(t\\right)={N}_{\\text{ss}},\\\\ {n}_{2}\\left(t\\right)={\\rm{e}}^{{\\lambda }_{\\text{ss}}\\left(1-r\\right)(t-{t}_{{\\rm{s}}})}\\end{array}$$<\/p>\n<p>\n                    (20)\n                <\/p>\n<p>As before, the SFS comprises variants that originated (1) before the positively selected mutation occurred, Si,1; (2) after this mutation occurred and happening in the mutant clone, Si,2; and (iii) after the mutation occurred but happening in normal stem cells, Si,3. Si,2 is the same as in the competition model and is modeled by equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ14\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>). Si,3 is now the superposition of Nss independent linear critical birth\u2013death processes and reads<\/p>\n<p>$${S}_{i,3}\\left(t\\right)={N}_{\\text{ss}}{\\int }_{\\!{t}_{\\text{S}}}^{t}{\\mu {\\lambda }_{\\text{ss}}P}_{\\text{ss},1,i}\\left(t-{t}^{{\\prime} }|{\\lambda }_{\\text{ss}}\\right)\\text{d}{t}^{{\\prime} }$$<\/p>\n<p>\n                    (21)\n                <\/p>\n<p>Likewise, Si,1 is now governed by genetic drift according to a critical birth\u2013death process in the founder cell population, in addition to selection of the mutant clone. It reads<\/p>\n<p>$$\\begin{array}{l}{S}_{i,1}\\left(t\\right)=\\\\\\left\\{\\begin{array}{c}\\mathop{\\sum }\\limits_{k=1}^{{N}_{\\text{ss}}}{S}_{k}\\left({t}_{\\text{s}}\\right)\\left[\\displaystyle\\frac{k}{{N}_{\\text{ss}}}{P}_{\\text{ss},k-1,i-C\\left(t\\right)}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}}\\right)\\right.\\\\\\left.+\\left(1-\\displaystyle\\frac{k}{{N}_{\\text{ss}}}\\right){P}_{\\text{ss},k,i}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}}\\right)\\right],i\\ge {n}_{2}\\left(t\\right),\\\\\\qquad\\qquad \\mathop{\\sum }\\limits_{k=1}^{{N}_{\\text{ss}}}{\\left(1-\\displaystyle\\frac{k}{{N}_{\\text{ss}}}\\right)}S_{k}\\left({t}_{\\text{s}}\\right){P}_{\\text{ss},k,i}\\left(t-{t}_{\\text{s}}|{\\lambda }_{\\text{ss}}\\right),i &lt; {n}_{2}(t).\\end{array}\\right.\\end{array}$$<\/p>\n<p>\n                    (22)\n                <\/p>\n<p>Modeling multiple selected clones in a homeostatic tissue. In the following, we generalize our model to account for the selection of multiple competing clones in a homeostatic tissue. We use \u03c4c to denote the time at which a particular clone c was born, where c\u2009=\u20091 identifies normal cells and 1\u2009&lt;\u2009c\u2009\u2264\u2009C identifies the C\u2009\u2212\u20091 selected clones. For each clone the variable vc reports the identity of its mother clone. We assume that all clones compete for a limited space with carrying capacity Nss. As with the one-clone model, we implement clonal selection as a reduction in the loss rates, while leaving division rates unchanged. Denoting the competition coefficient between clone c and other clones in the tissue with \\({\\rho }_{c\\bullet }\\), we model the expected number of cells in clone c, nc, with a system of ordinary differential equations,<\/p>\n<p>$$\\frac{{\\text{d}n}_{c}}{\\text{d}t}={\\lambda }_{\\text{ss}}{n}_{c}\\left(1-\\mathop{\\sum }\\limits_{j;1\\le\\; j\\le C}{\\rho }_{{jc}}\\frac{{n}_{j}}{{N}_{\\text{ss}}}\\right)$$<\/p>\n<p>\n                    (23)\n                <\/p>\n<p>with conditions<\/p>\n<p>$$\\begin{array}{c}{n}_{c}\\left({\\tau }_{c}\\right)=1,\\\\ {n}_{c}\\left({t &lt; \\tau }_{c}\\right)=0,\\\\ {n}_{{v}_{c}}\\left({\\tau }_{c}\\right)={n}_{{v}_{c}}\\left({\\tau }_{c}-{\\rm{d}t}\\right)-1.\\end{array}$$<\/p>\n<p>\n                    (24)\n                <\/p>\n<p>The competition matrix \u03c1 is defined by the selective advantages, and, requiring \u03a3jnj\u2009=\u2009Nss at all times, is given by<\/p>\n<p>$$\\rho =\\left(\\begin{array}{cccc}1 &amp; 2-{r}_{2} &amp; 2-{r}_{3} &amp; \\ldots \\\\ {r}_{2} &amp; 1 &amp; 2-{r}_{3}\/{r}_{2} &amp; \\ldots \\\\ {r}_{3} &amp; {r}_{3}\/{r}_{2} &amp; 1 &amp; \\ldots \\\\ \\ldots &amp; \\ldots &amp; \\ldots &amp; \\ldots \\end{array}\\right)$$<\/p>\n<p>where rc models a reduction of cell loss in clone c at \u03c4c and takes values 0\u2009&lt;\u2009rc\u2009&lt;\u20091. Ultimately, we are interested in Si(t), the expected number of variants that are found in a clone of size i at the time of measurement, t. As in the one-clone model, Si(t) has contributions from variants acquired in normal cells and variants acquired in the selected clones. In the following, we denote the SFS of variants acquired in a particular clone with Si,c(t). Hence,<\/p>\n<p>$${S}_{i}\\left(t\\right)=\\mathop{\\sum }\\limits_{c=1}^{C}{S}_{i,c}\\left(t\\right)$$<\/p>\n<p>\n                    (25)\n                <\/p>\n<p>To evaluate Si,c(t), we denote with \\({\\phi }_{c}\\) the set of subclones generated by a mutation in clone c, in the following termed \u2018daughters\u2019 of clone c, and with \\(\\varphi \\subset {\\phi }_{c}\\) the set of all possible combinations of \\({\\phi }_{c}\\). Finally, the set \u03a8c denotes the entire progeny of clone c, including daughters, granddaughters and so on. Consider a variant that was acquired in clone c and is present in k cells at time t\u2032. This variant will, on average, be inherited to the daughter clone \\(d,d\\in {\\phi }_{c}\\), with probability<\/p>\n<p>$$\\kappa \\left(c,d,k,t{\\prime} \\right)=H({\\tau }_{d}-{t}^{{\\prime} })\\frac{k}{{n}_{c}(t{\\prime} )}$$<\/p>\n<p>\n                    (26)\n                <\/p>\n<p>where H is the heavyside step function and \u03c4d is the birth date of clone d. Considering a particular subset of daughters of clone c, \\({\\varphi }_{j}\\in \\varphi\\), the probability of inheriting the variant to all daughters in this subset is, on average,<\/p>\n<p>$$\\begin{array}{l}{{{K}}}\\left(c,{\\varphi }_{j},k,t{\\prime} \\right)=\\overbrace{{\\mathop{\\prod}\\nolimits_{l\\in {\\varphi }_{j},{\\varphi }_{j}\\in \\varphi }}\\kappa \\left(c,l,k,{t}^{{\\prime} }\\right)}^{{\\textrm{presence\\; in\\; the\\; subset\\; of\\; daughters\\;}}{\\varphi }_{j}\\in \\varphi }\\\\\\times \\underbrace{{\\mathop{\\prod}\\nolimits _{l\\in \\left\\{{\\phi }_{c}{\\rm{\\backslash }}{\\varphi }_{j}\\right\\},{\\varphi }_{j}\\in \\varphi }}\\left(1-\\kappa \\left(c,l,k,{t}^{{\\prime} }\\right)\\right)}_{{\\textrm{absence\\; in\\; the\\; remaining\\; daughters\\;}}}\\end{array}$$<\/p>\n<p>\n                    (27)\n                <\/p>\n<p>We compute Si,c(t) by considering all possible combinations of daughters that may inherit variants acquired in clone c. This yields<\/p>\n<p>$$\\begin{array}{l}{S}_{i,c}\\left(t\\right)\\\\=\\overbrace{{\\displaystyle\\int}_{\\!{\\tau }_{c}}^{t}\\,{\\stackrel{\\textrm{acquisition\\; of\\; variants}}{\\mu {\\lambda }_{\\text{ss}}{n}_{c}(t{\\prime} )}}\\,\\underbrace{\\sum _{{\\varphi }_{j}\\in \\varphi }{{{\\rm K}}}\\left(c,{\\varphi }_{j},1,t{\\prime} \\right){P}_{c}\\left(1,{\\stackrel{\\textrm{drift\\; within\\; clone\\;}c}{i-\\sum _{k\\in {\\varphi }_{j}}{n}_{k}\\left(t\\right)}},{t}^{{\\prime} },t\\right)}_{{\\textrm{all\\; possible\\; combinations\\; of\\; daugthers}}}\\text{d}{t}^{{\\prime} }+}^{{\\textrm{variants\\; acquired\\; in\\; clone\\;}}c\\;{\\textrm{and\\; inherited\\; to\\; at\\; least\\;}}1\\;{\\textrm{daughter\\; of\\; clone\\;}}c}\\\\ \\underbrace{{\\displaystyle\\int}_{\\!{\\tau}_{c}}^{t}\\overbrace{\\mu {\\lambda }_{\\text{ss}}{n}_{c}(t{\\prime} )}^{{\\textrm{acquisition\\; of\\; variants}}}\\left(1-\\sum _{{\\varphi }_{j}\\in \\varphi }{{{\\rm K}}}\\left(c,{\\varphi }_{j},1,t{\\prime} \\right)\\right){P}_{c}(1,i,{t}^{{\\prime} },t)\\text{d}{t}^{{\\prime} }}_{{\\textrm{variants\\; acquired\\; in\\; clone\\;}}c\\;{\\textrm{and\\; not\\; inherited\\; to\\; any\\; daughter\\; of\\; clone\\;}}c}\\end{array}$$<\/p>\n<p>\n                    (28)\n                <\/p>\n<p>where Pc(1,i,t\u2032,t) is the probability that a variant acquired in clone c has drifted to i cells within a time span t\u2009\u2212\u2009t\u2032. Specifically, Pc(1,i,t\u2032,t) is defined by nonlinear birth\u2013death processes, because the division and loss rates in clone c change with time, subject to the dynamics in equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ23\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>). In analogy to the one-clone model, we approximate this nonlinearity by modeling Si,c(t) with linear birth\u2013death processes in a step-wise fashion, considering time intervals between clonal birth dates, \u03c4 turning points in n(t) (as determined by numeric evaluation of equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ23\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>)) and the end point as time points of interest. Specifically, we modeled the drift of somatic variants in each time interval using linear birth\u2013death processes, which we parametrized with an effective loss rate of cells in clone c, \\({\\delta }_{\\text{eff},c,{t}_{a}\\le t &lt; {t}_{b}},\\) and which we evaluated for an effective time span \\({\\varDelta }_{\\text{eff},{c,t}_{a}\\le t &lt; {t}_{b}}\\), where ta and tb denote the start and end point of the interval. To define \\({\\delta }_{\\text{eff},c,{t}_{a}\\le t &lt; {t}_{b}}\\) and \\({\\varDelta }_{c,\\text{eff},{t}_{a}\\le t &lt; {t}_{b}}\\), we distinguished the case when the number of cells in clone c increased (nc(tb)\u2009&gt;\u2009nc(ta)) from the case when it decreased (nc(tb)\u2009&lt;\u2009nc(ta)) during the time interval (ta,tb). If nc(tb)\u2009&gt;\u2009nc(ta), we defined \\({\\delta }_{\\text{eff},c,{(t}_{a},{t}_{b})}={r}_{c}{\\lambda }_{\\text{ss}}\\) and, to avoid overshooting the actual clone size, \\({\\varDelta }_{{\\rm{eff}},c,(t_{a},{t}_{b})}=\\frac{\\log {n}_{c}({t}_{b})\/{n}_{c}({t}_{a})}{{\\lambda }_{\\text{ss}}(1-{r}_{c})}\\). By contrast, if nc(tb)\u2009&lt;\u2009nc(ta), we parametrized \\({\\delta }_{\\text{eff},c{(,t}_{a},{t}_{b})}\\) such that the expected decline in a linear birth\u2013death process equals the number of cells lost by competition dynamics. Hence, \\({\\delta }_{{\\text{eff}},c({t}_{a},{t}_{b})}\\) is defined by<\/p>\n<p>$$\\begin{array}{l}{\\displaystyle\\int }_{\\!{{\\rm{t}}}_{{\\rm{a}}}}^{{{\\rm{t}}}_{{\\rm{b}}}}{\\delta }_{\\text{eff},c,(t_{a},{t}_{b})}{n}_{{{c}}}({t}_{a})\\exp (({{\\rm{\\lambda }}}_{\\text{ss}}-{\\delta }_{\\text{eff},c,(t_{a},{t}_{b})}){\\rm{t}}{\\prime} )\\text{d}t{\\prime} \\\\={\\displaystyle\\int }_{\\!{{\\rm{t}}}_{{\\rm{a}}}}^{{{\\rm{t}}}_{{\\rm{b}}}}{{\\rm{\\lambda }}}_{\\text{ss}}{n}_{{{c}}}({t}^{{\\prime} })\\mathop{\\sum }\\limits_{1\\le\\; j\\le C}{{\\rho }_{jc}}\\frac{{n}_{j}({t}^{{\\prime} })}{{N}_{{\\rm{ss}}}}\\text{d}t{\\prime}\\end{array}$$<\/p>\n<p>\n                    (29)\n                <\/p>\n<p>where the right-hand side gives the number of death events in clone c according to the competition dynamics in equation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"equation anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Equ23\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a>) and the left-hand side gives the number of death events according to a linear birth\u2013death process. Moreover, if nc(tb)\u2009&lt;\u2009nc(ta), we defined \\({\\varDelta }_{c,{\\rm{eff}},({t}_{a},{t}_{b})}={t}_{b}-{t}_{a}\\). We then approximated Si,c(t) by recursively evaluating<\/p>\n<p>$$\\begin{array}{l}{S}_{i,c}\\left({t}_{b}\\right)=\\mathop{\\sum }\\limits_{k=1}^{{N}_{\\text{ss}}}{S}_{k,c}\\left({t}_{a}\\right)\\\\\\left[\\overbrace{\\sum _{{\\varphi }_{j}\\in \\varphi }{{{K}}}\\left(c,{\\varphi }_{j},k,{t}_{a}\\right){P}_{\\exp ,k,i-\\sum _{l\\in {\\varphi }_{j}}{n}_{l}\\left(t\\right)}\\left({\\varDelta }_{{\\rm{eff}},c,(t_{a}{,t}_{b})}|{{\\lambda }_{\\text{ss}},\\delta }_{\\text{eff},c,(t_{a},{t}_{b})}\\right)}^{A}\\right.\\\\\\left.+\\overbrace{\\left(1-\\sum _{{\\varphi }_{j}\\in \\varphi }{{{K}}}\\left(c,{\\varphi }_{j}k,{t}_{a}\\right)\\right){P}_{\\exp ,k,i}\\left({\\varDelta }_{{\\rm{eff}},c,({t}_{a},{t}_{b})}|{{\\lambda }_{\\text{ss}},\\delta }_{\\text{eff},c,(t_{a},{t}_{b})}\\right)}^{B}\\right]\\\\+\\overbrace{{n}_{c}\\left({t}_{a}\\right){\\int }_{{t}_{a}}^{{t}_{b}}\\mu {\\lambda }_{\\text{ss}}{\\rm{e}}^{({\\lambda }_{\\text{ss}}-{\\delta }_{\\text{eff},c,({t}_{a},{t}_{b})})({\\varDelta }_{c,\\text{eff},({t}_{a},{t}_{b})})}{P}_{\\exp ,1,i}\\left({\\varDelta }_{{\\rm{eff}},c,(t_{a},{t}_{b})}|{{\\lambda }_{\\text{ss}},\\delta }_{\\text{eff},c,(t_{a},{t}_{b})}\\right)}^{C}\\end{array}$$<\/p>\n<p>\n                    (30)\n                <\/p>\n<p>until tb\u2009=\u2009t. Here, terms (A) and (B) compute the drift of variants acquired before ta, of which some (A) will be inherited to subclones of clone c, whereas others (B) will be found exclusively in clone c. Finally, term (C) computes the drift of variants newly acquired in the time interval (ta,tb).<\/p>\n<p>                  Parameter estimation<\/p>\n<p>Model fits were obtained for each donor separately. In a first step, clonal selection was distinguished from genetic drift by applying the one-clone model of SCIFER to the data. For blood samples classified as selected and sequenced at high resolution (\u2265150\u00d7 WGS), subsequent refinement was achieved by applying the two-clone model in a second step. Here, both possible topologies between two selected clones (linear and branched evolution) were tested, using prior distributions that were informed by the initial one-clone model fits (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>; note that constraining the prior distributions based on the one-clone model fits facilitates the detection of additional subclones). All model fits were performed using ABC based on sequential Monte Carlo as implemented in pyABC<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Klinger, E., Rickert, D. &amp; Hasenauer, J. pyABC: distributed, likelihood-free inference. Bioinformatics 34, 3591&#x2013;3593 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR82\" id=\"ref-link-section-d19200088e16685\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> (posterior sample size of 1,000 and termination criterion \u03b5min\u2009=\u20090.05). Briefly, ABC samples parameter sets from user-defined prior distributions (Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>) and simulates the expected VAF distribution for each parameter set. The simulated VAF distributions are compared with the measured ones and the 1,000 parameter sets by minimal distance between measured and simulated data are selected. In each of the following iterations, the parameter set is expanded by adding random noise to the current parameter set, and the distance between model and data is reassessed. ABC is terminated once the distance between model and data is smaller than a defined threshold, \u03b5min, yielding an estimate for the posterior distributions of the model parameters. We performed the following steps in each ABC iteration:<\/p>\n<p>                      (1)<\/p>\n<p>[Experimental data]: Determine the cumulative VAF distribution at time t for the experimental data. For 90\u00d7 bulk WGS data from human BM or peripheral blood samples, include variants if they are supported by at least 3 reads, absent in the germline control sample, and if the locus is covered by at least 10 and at most 300 reads. For 270\u00d7 bulk WGS data from human CD34+ HSPCs, include variants if they are supported by at least 3 reads. For pseudo-bulk WGS data, include all variants. For bulk WGS data from human brain, include all tier1\u2013tier2 variants from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 17\" title=\"Bae, T. et al. Analysis of somatic mutations in 131 human brains reveals aging-associated hypermutability. Science 377, 511&#x2013;517 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR17\" id=\"ref-link-section-d19200088e16719\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>. Compute the cumulative number of VAFs \u2265f, defined as \\({M}_{\\text{experimental},\\;f}={\\sum }_{\\text{VAF}=f}^{1}{\\;S}_{\\text{experimental},\\text{VAF}},\\) where Sexperimental denotes the measured SFS and Sexperimental,VAF is the number of variants with a particular VAF. We evaluated Mexperimental,f for bins with width 0.01, spanning 0.05\u2009\u2264\u2009f\u2009\u2264\u20091 for bulk WGS data with an average coverage &lt;150\u00d7, 0.02\u2009\u2264\u2009f\u2009\u2264\u20091 for bulk WGS data with an average coverage \u2265150\u00d7, and 0.01\u2009\u2264\u2009f\u2009\u2264\u20091 for pseudo-bulk data.<\/p>\n<p>                      (2)<\/p>\n<p>[Model]: Sample parameters from their prior distributions (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>, one-clone model; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>, two-clone model).<\/p>\n<p>                      (3)<\/p>\n<p>[Model]: Simulate the expected cumulative VAF distribution, Msim,f(tdata), where tdata is the age of the patient and f the minimal VAF (for numerical implementation of the model see Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). The bins are as with the experimental data. Depending on the following criterion, based on the prior parameter sample, the cumulative VAF histogram is simulated with either the neutral model or the selection model: if a selected clone could grow above the detection limit with the prior parameter sample, the selection model is used; otherwise the neutral model is used. Formally, if \\(\\frac{1}{2}{\\rm{e}}^{{\\lambda }_{{\\rm{SS}}}^{{\\rm{prior}}}(1-{r}^{{\\rm{prior}}})({t}_{\\text{data}}-{t}_{{\\rm{S}}}^{{\\rm{prior}}})}\\ge \\gamma {N}_{{\\rm{SS}}}^{{\\rm{prior}}}\\) (one-clone model) or if any \\(\\frac{1}{2}{n}_{c,c &gt; 1}\\left({t}_{{\\rm{data}}}|{\\lambda }_{{\\rm{SS}}}^{{\\rm{prior}}},{{\\tau }^{{\\rm{prior}}},r}^{{\\rm{prior}}},\\phi ,\\psi {,N}_{{\\rm{SS}}}^{{\\rm{prior}}}\\right)\\ge \\gamma {N}_{{\\rm{SS}}}^{{\\rm{prior}}}\\) (two-clone model) then the selection model is used. The detection limit is set to \u0263\u2009=\u20090.025 for 90\u00d7 and 30\u00d7 WGS data, and \u0263\u2009=\u20090.005 for pseudo-bulk and 270\u00d7 bulk WGS data. Importantly, note that the selection model can return a posterior corresponding to neutral evolution.<\/p>\n<p>                      (4)<\/p>\n<p>[Model]: Add the number of variants present in the founder cell of the hematopoietic system, \\({\\varDelta }_{{\\rm{clonal}}}^{{\\rm{prior}}}\\), to Msim,0.5 (corresponding to mutations acquired during early development).<\/p>\n<p>                      (5)<\/p>\n<p>[Model]: Simulate experimental error of sequencing. To this end, generate the expected SFS for the jth frequency bin, Ssim(fj)\u2009=\u2009Msim(fj) \u2212\u2009Msim(fj+1). For bulk WGS data, sample for each simulated variant a sequencing coverage \u03d1 from a Poisson distribution with mean \\(\\hat{\\vartheta }\\) corresponding to the average sequencing depth. Thereafter, sample VAFs for each variant according to \\(\\frac{B\\left(f,\\vartheta \\right)}{\\vartheta}\\), where B denotes the binomial distribution and f is the true VAF in the tissue. Discard variants supported by less than 3 reads. For pseudo-bulk WGS data, sample VAFs for each variant according to \\(\\frac{B\\left(2f,{n}_{\\text{cells}}\\right)}{(2{n}_{\\text{cells}})}\\), where \\({n}_{\\text{cells}}\\) is the number of sequenced single-cell clones. Compute the sampled cumulative VAF distribution, Msim,f,sampled.<\/p>\n<p>                      (6)<\/p>\n<p>[Model versus Experimental data]: Determine the distance function for ABC<\/p>\n<p>$$d=\\mathop{\\sum }\\limits_{f}{\\left({M}_{\\text{sim},\\;f,\\text{sampled}}-{M}_{\\text{experimental},\\;f}\\right)}^{2}$$<\/p>\n<p>Classification of cases as neutrally evolving or as selected<\/p>\n<p>We classified pseudo-bulks and samples sequenced at &lt;150\u00d7 as selected if at least 15% of the posterior samples report a selected clone size \u22650.1. These thresholds were defined based on in silico generated test data, and correspond to the WGS depth of 90\u00d7 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). WGS data with \u2265150\u00d7 coverage allow for higher resolution and hence, we classified cases as selected if at least 15% of the posterior samples report a selected clone size \u22650.04 according to the one-clone model, and as harboring two selected clones if at least 15% of the posterior samples report a size \u22650.04 for both selected clones. Upon sample classification, we computed the 80% highest density intervals for each parameter on the parameter subsets supporting neutral evolution or clonal selection, respectively.<\/p>\n<p>Simulation of phylogenetic trees and in silico evaluation of model performance<\/p>\n<p>We validated the population genetics model with simulated data, generated according to stochastic birth\u2013death processes (see Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> for a description of the simulations and <a href=\"https:\/\/github.com\/VerenaK90\/SCIFER\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/VerenaK90\/SCIFER<\/a> for their computational implementation).<\/p>\n<p>Simulations used to evaluate model performance (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>) were run for stem cells only. Simulations of neutral evolution were parametrized with Nss,S\u2009=\u200925,000, \u03bc\u2009=\u20091, \u03bbss,S\u2009=\u200910 per year, \u03b4ss,S\u2009=\u200910 per year and with \u03c4\u2009=\u2009250 (that is, summarizing 1% of the reactions occurring in 25,000 stem cells). Simulations of clonal selection were parametrized with Nss,S\u2009=\u200925,000, \u03bc\u2009=\u20091, \u03bbss,S\u2009=\u200910 per year, \u03b4ss,S\u2009=\u200910 per year, ts\u2009=\u200920\u2009years, s\u2009=\u20090.02 and with \u03c4\u2009=\u20092,500.<\/p>\n<p>Simulations of neutral evolution used to assess the effect of differentiation into a single progenitor cell population on the VAF histogram were parametrized with Nss,S\u2009=\u20091,000, \u03bc\u2009=\u20091, \u03bbss,S\u2009=\u20091 per year, \u03b4ss,S\u2009=\u20091 per year, Nss,P\u2009=\u20092,500, \u03bbss,P\u2009=\u20094.6 per year, \u03b4ss,S\u2009=\u20095 per year and \u03c4\u2009=\u20092,500.<\/p>\n<p>Simulations of neutral evolution in a heterogenous tissue where stem cells differentiate into two progenitor cell types during development, but continue to produce only one of them during adulthood were parametrized with Nss,S\u2009=\u20091,000, \u03bc\u2009=\u20091, \u03bbss,S\u2009=\u20091 per year, \\({\\delta }_{\\text{ss},\\text{S} &gt; {P}_{1}}\\)\u2009=\u20091 per year, \\({\\delta }_{\\text{ss},\\text{S} &gt; {P}_{2}}\\)\u2009=\u20090 per year, \\({N}_{\\text{ss},{P}_{1}}\\)\u2009=\u20096,000 \\({\\lambda }_{\\text{ss},{P}_{1}}\\)\u2009=\u20093 per year, \\({\\delta }_{\\text{ss},{P}_{1}}\\)\u2009=\u20093.167 per year, \\({N}_{\\text{ss},{P}_{2}}\\)\u2009=\u20094,000, \\({\\lambda }_{\\text{ss},{P}_{2}}\\)\u2009=\u20090 per year, \\({\\delta }_{\\text{ss},{P}_{2}}\\)\u2009=\u20090 per year and \u03c4\u2009=\u20091,000.<\/p>\n<p>Parameter inference<\/p>\n<p>To infer the dynamic stem cell parameters (\u03bc, \u03b4exp, \u03bbss,S, Nss,S, ts, r) (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>), we subsampled 10,000 cells from the simulated trees at time points specified in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> and computed the simulated VAF distribution from subsampled trees. To account for technical noise, we simulated sequencing by sampling for each variant a sequencing coverage \\(\\vartheta \\propto {Pois}(\\hat{\\vartheta })\\), where \u03d1 is the average sequencing coverage and thereafter sampling mutant reads according to B(VAF,\u03d1), where B is the binomial distribution. We simulated VAFs for average sequencing coverages of 30\u00d7, 90\u00d7 and 270\u00d7, and fitted the population genetics model to the simulated bulk WGS data as described above for the real data.<\/p>\n<p>Sensitivity and specificity of detecting selected clones<\/p>\n<p>To evaluate the sensitivity and specificity of the out model, we analyzed the posterior probability of clonal selection. In accordance with the minimal clone sizes used in the inference setup (see above), we computed the probability of clonal selection as \\(P\\left({\\rm{selection}}\\right)=\\frac{{\\sum }_{i}\\frac{{n}_{2}({t}_{\\text{data}}|{\\theta }_{i})}{{N}_{\\text{ss},i}}\\ge 0.05}{{\\sum }_{i}1}\\) for 30\u00d7 and 90\u00d7 sequencing depths, and as \\(P\\left({\\rm{selection}}\\right)=\\frac{{\\sum }_{i}\\frac{{n}_{2}({t}_{\\text{data}}|{\\theta }_{i})}{{N}_{\\text{ss},i}}\\ge 0.01}{{\\sum }_{i}1}\\) for 270\u00d7 sequencing depth, where \\({n}_{2}({t}_{\\text{data}}|{\\theta }_{i})\\) is the size of the selected clone at the patient age, tdata, given the ith parameter sample, \u03b8i, and Nss,i is the ith estimate of the stem cell number. Among all parameter sets reporting a clone size \u22650.05 (30\u00d7 and 90\u00d7 coverage) or \u22650.01 (270\u00d7 coverage), we determined median clone size and, if they were at least 5% in size, rounded it with 5% accuracy, or else rounded them with 1% accuracy. Thereafter, we computed the sensitivity and specificity of our approach, by classifying cases as selected if P(selection) \u2265 \u03b2, varying the threshold \u03b2 between 1% and 100%. For each of the true clone sizes (1%, 2%, 5%, 10%, 15%, 20%, 25%, 50%, 75%; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>), we computed the number of true positives as the number of cases where the actual clone size was correctly inferred (we classified inferred clone sizes between 0.5 and 1.5 of the actual clone size as correct). Conversely, we computed the number of false positives as the number of cases in which SCIFER erroneously reported a particular clone size, albeit the actual clone size was zero.<\/p>\n<p>The difference between true positives and false positives was maximal for \u03b2\u2009=\u200915%. At this threshold, clones of size \u22655% VAF were reliably inferred (90\u00d7).<\/p>\n<p>Statistics and reproducibility<\/p>\n<p>Bayesian parameter inference and statistical analyses were conducted with pyABC<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Klinger, E., Rickert, D. &amp; Hasenauer, J. pyABC: distributed, likelihood-free inference. Bioinformatics 34, 3591&#x2013;3593 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR82\" id=\"ref-link-section-d19200088e18752\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> v.0.12.6, using python v.3.10.1 and R (v.4.2.0 and v.4.2.1). We used the following R packages: ape<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 83\" title=\"Paradis, E. &amp; Schliep, K. ape 5.0: an environment for modern phylogenetics and evolutionary analyses in R. Bioinformatics 35, 526&#x2013;528 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR83\" id=\"ref-link-section-d19200088e18756\" rel=\"nofollow noopener\" target=\"_blank\">83<\/a> v.5.6-2, phytools<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 84\" title=\"Revell, L. J. phytools: an R package for phylogenetic comparative biology (and other things). Methods Ecol. Evol. 3, 217&#x2013;223 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR84\" id=\"ref-link-section-d19200088e18760\" rel=\"nofollow noopener\" target=\"_blank\">84<\/a> v.1.2-0, phangorn<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 85\" title=\"Schliep, K., Potts, A. A., Morrison, D. A. &amp; Grimm, G. W. Intertwining phylogenetic trees and networks. Methods Ecol. Evol. 8, 1212&#x2013;1220 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR85\" id=\"ref-link-section-d19200088e18764\" rel=\"nofollow noopener\" target=\"_blank\">85<\/a> v.2.10.0, castor<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 86\" title=\"Louca, S. &amp; Doebeli, M. Efficient comparative phylogenetics on large trees. Bioinformatics 34, 1053&#x2013;1055 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR86\" id=\"ref-link-section-d19200088e18768\" rel=\"nofollow noopener\" target=\"_blank\">86<\/a> v.1.7.5, TreeTools v.1.8.0, deSolve<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 87\" title=\"Soetaert, K., Petzoldt, T. &amp; Setzer, R. W. Solving differential equations in R: package deSolve. J. Stat. Softw. 33, 1&#x2013;25 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR87\" id=\"ref-link-section-d19200088e18773\" rel=\"nofollow noopener\" target=\"_blank\">87<\/a> v.1.33, openxlsx v.4.2.5, cdata v.1.2.0, ggpubr v.0.4.0, RRphylo<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 88\" title=\"Castiglione, S. et al. A new method for testing evolutionary rate variation and shifts in phenotypic evolution. Methods Ecol. Evol. 9, 974&#x2013;983 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR88\" id=\"ref-link-section-d19200088e18777\" rel=\"nofollow noopener\" target=\"_blank\">88<\/a> v.2.7.0, ggplot2 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 89\" title=\"Villanueva, R.A.M. &amp; Chen, Z.J. ggplot2: Elegant Graphics for Data Analysis (Taylor &amp; Francis, 2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR89\" id=\"ref-link-section-d19200088e18781\" rel=\"nofollow noopener\" target=\"_blank\">89<\/a>) v.3.4.2, cgwtools v.3.3, ggVennDiagram v.1.2.2, ggbeeswarm v.0.6.0, ggsci v.2.9, Hmisc v.4.7.1, lemon v.0.4.5, data.table v.1.14.2, RColorBrewer v.1.1.3, ggridges v.0.5.4, doParallel v.1.0.17, foreach v.1.5.2, parallel v.2.1, wesanderson v.0.3.6, bedr v.1.0.7, ggformula v.0.10.2, HDInterval v.0.2.2, reshape2 v.1.4.4 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 90\" title=\"Wickham, H. Reshaping data with the reshape package. J. Stat. Softw. 21, 1&#x2013;20 (2007).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#ref-CR90\" id=\"ref-link-section-d19200088e18785\" rel=\"nofollow noopener\" target=\"_blank\">90<\/a>), dplyr v.1.0.9 and scales v.1.2.1. Flow cytometry data was analyzed with FlowJo v.10.8.1.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02217-y#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Ethical approval Patient samples were collected with informed consent under the Mechanisms of Age-Related Clonal Haematopoiesis (MARCH) Study.&hellip;\n","protected":false},"author":2,"featured_media":19168,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25],"tags":[7736,6634,7003,64,63,5562,1617,5563,7002,1325,336,7001,20186,128,20187],"class_list":["post-19167","post","type-post","status-publish","format-standard","has-post-thumbnail","category-genetics","tag-ageing","tag-agriculture","tag-animal-genetics-and-genomics","tag-au","tag-australia","tag-biomedicine","tag-cancer","tag-cancer-research","tag-gene-function","tag-general","tag-genetics","tag-human-genetics","tag-population-genetics","tag-science","tag-systems-biology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/19167","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=19167"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/19167\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/19168"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=19167"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=19167"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=19167"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}