{"id":217141,"date":"2025-10-16T11:15:20","date_gmt":"2025-10-16T11:15:20","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/217141\/"},"modified":"2025-10-16T11:15:20","modified_gmt":"2025-10-16T11:15:20","slug":"sperm-sequencing-reveals-extensive-positive-selection-in-the-male-germline","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/217141\/","title":{"rendered":"Sperm sequencing reveals extensive positive selection in the male germline"},"content":{"rendered":"<p>Ethics<\/p>\n<p>This study was carried out under TwinsUK BioBank ethics, approved by the North West\u2013Liverpool Central Research Ethics Committee (REC reference 19\/NW\/0187), IRAS ID 258513 and earlier approvals granted to TwinsUK by the St Thomas\u2019 Hospital Research Ethics Committee, later the London\u2013Westminster Research Ethics Committee (REC reference EC04\/015).<\/p>\n<p>Sample collection<\/p>\n<p>Semen samples were collected or obtained from archival samples with informed consent from 75 research participants in the TwinsUK cohort<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Verdi, S. et al. TwinsUK: The UK Adult Twin Registry Update. Twin Res. Hum. Genet. 22, 523&#x2013;529 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR25\" id=\"ref-link-section-d7063975e1837\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>. Archival whole-blood samples were also obtained from 67 of those men from the TwinsUK BioBank. A total of 104 semen samples spanned an age range of 24\u201375\u2009years and included 29 men with 2\u2009time points separated by a mean of 12.1\u2009years (range of 12\u201313\u2009years) and the remaining 46 men with a single time point. A total of 133 blood samples were collected at an age range of 22\u201383\u2009years from men with 1\u20134\u2009time points. The mean interval between blood time points was 8.1\u2009years (range of 1\u201313\u2009years). In the cohort, there were a total of nine monozygotic twins and three dizygotic twin pairs. Counts of samples, time points and twin pairs that were successfully sequenced and passed analysis quality control thresholds are summarized in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>.<\/p>\n<p>Metadata<\/p>\n<p>Self-reported age, height, weight, ethnicity, twin zygosity, smoking and alcohol consumption were obtained from questionnaires provided by TwinsUK and periodically taken. All individuals who provided ethnicity information indicated white. BMI was calculated as weight\/height2. A smoking pack-years was defined as 365 packs of cigarettes and a total pack-year was calculated using the highest estimate across all questionnaires from cigarettes per day or week and total years smoked. Alcohol drink-years was calculated using average weekly alcohol consumption extrapolated to the duration of adult life before sampling (age\u2009\u2013\u200918).<\/p>\n<p>DNA extraction<\/p>\n<p>DNA was extracted from sperm samples using a Qiagen QIAamp DNA Blood Mini kit. Isolation of genomic DNA from sperm protocol\u20091 (QA03 Jul-10) was followed, but with the exceptions of substituting DTT in place of \u03b2-mercaptoethanol for buffer\u20092 and substituting buffer EB in place of buffer AE for the elution of DNA.<\/p>\n<p>DNA was extracted from whole blood using a Gentra Puregene Blood kit using the protocol for 10\u2009ml of compromised whole blood from the Gentra Puregene Handbook (v.06\/2011).<\/p>\n<p>Targeted gene panels<\/p>\n<p>Three separate Twist Bioscience gene panels were used for targeted NanoSeq sequencing in this study: (1) a custom pilot panel of 210 genes; (2) a similar but extended custom panel of 263 genes (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>); (3) and a default exome-wide panel of 18,800 genes. The two custom gene panels were highly similar, with the extended panel being almost exclusively regions added to the pilot panel. From the 84 samples that underwent targeted sequencing, 13 were sequenced using a pilot panel of 210 canonical cancer or somatic driver genes, and all 84 were sequenced using the extended panel of 263 genes. Sequencing coverage and mutations were merged from samples sequenced on both targeted panels and they were treated as \u2018targeted\u2019 samples in the Article. The custom panels were designed by gathering sets of published lists of genes implicated as drivers in cancers<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Gao, Y.-B. et al. Genetic landscape of esophageal squamous cell carcinoma. Nat. Genet. 46, 1097&#x2013;1102 (2014).\" href=\"#ref-CR61\" id=\"ref-link-section-d7063975e1877\">61<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Waddell, N. et al. Whole genomes redefine the mutational landscape of pancreatic cancer. Nature 518, 495&#x2013;501 (2015).\" href=\"#ref-CR62\" id=\"ref-link-section-d7063975e1877_1\">62<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Aaltonen, L. A. et al. Pan-cancer analysis of whole genomes. Nature 578, 82&#x2013;93 (2020).\" href=\"#ref-CR63\" id=\"ref-link-section-d7063975e1877_2\">63<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 64\" title=\"Nguyen, B. et al. Genomic characterization of metastatic patterns from prospective clinical sequencing of 25,000 patients. Cell 185, 563&#x2013;575 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR64\" id=\"ref-link-section-d7063975e1880\" rel=\"nofollow noopener\" target=\"_blank\">64<\/a> and somatic tissues<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 6\" title=\"Martincorena, I. et al. High burden and pervasive positive selection of somatic mutations in normal human skin. Science 348, 880&#x2013;886 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR6\" id=\"ref-link-section-d7063975e1884\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 65\" title=\"Kakiuchi, N. &amp; Ogawa, S. Clonal expansion in non-cancer tissues. Nat. Rev. Cancer 21, 239&#x2013;256 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR65\" id=\"ref-link-section-d7063975e1887\" rel=\"nofollow noopener\" target=\"_blank\">65<\/a> as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Lawson, A. R. J. et al. Somatic mutation and selection at population scale. Nature &#010;                https:\/\/doi.org\/10.1038\/s41586-025-09584-w&#010;                &#010;               (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR5\" id=\"ref-link-section-d7063975e1891\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>.<\/p>\n<p>Sequencing and library preparation<\/p>\n<p>Restriction-enzyme whole-genome NanoSeq libraries were prepared as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e1903\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> and subjected to whole-genome sequencing (WGS) at target 20\u201330\u00d7 coverage on NovaSeq (Illumina) platforms to generate 150-bp paired-end reads with 9\u201310 samples per lane. Standard WGS of blood (31.7\u00d7 median coverage) was used to generate matched-normal libraries for both restriction-enzyme NanoSeq blood and sperm samples.<\/p>\n<p>Targeted and exome NanoSeq libraries were prepared by sonication and one to two rounds of pull down of target sequences. They were then PCR amplified and sequenced with NovaSeq (Illumina) platforms to generate 150-bp paired-end reads with 7\u20138 samples per lane for the targeted panel and 2 lanes per sample for the exome panel. These steps are described in detail in \u2018Sonication NanoSeq, Library amplification and sequencing, and Hybridization Capture\u2019 of supplementary note\u20091 of ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 5\" title=\"Lawson, A. R. J. et al. Somatic mutation and selection at population scale. Nature &#010;                https:\/\/doi.org\/10.1038\/s41586-025-09584-w&#010;                &#010;               (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR5\" id=\"ref-link-section-d7063975e1910\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>.<\/p>\n<p>Base calling and filtering<\/p>\n<p>All samples were processed using a Nextflow implementation of the NanoSeq calling pipeline (<a href=\"https:\/\/github.com\/cancerit\/NanoSeq\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/cancerit\/NanoSeq<\/a>). BWA-MEM<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 66\" title=\"Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. Preprint at &#010;                https:\/\/doi.org\/10.48550\/arXiv.1303.3997&#010;                &#010;               (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR66\" id=\"ref-link-section-d7063975e1929\" rel=\"nofollow noopener\" target=\"_blank\">66<\/a> was used to align all sequences to the human reference genome (NCBI build37). Whole-genome NanoSeq samples were called with their matched WGS normal and default parameters except for var_b (minimum matched normal reads per strand) of 5 as needed for WGS normals, cov_Q (minimum mapQ to include a duplex read) of 15 and var_n (maximum number of mismatches) of 2.<\/p>\n<p>For targeted and exome samples, we leveraged the high sequencing depth and high polyclonality to exclude variants with VAF\u2009&gt;\u200910% instead of using a matched normal. Default parameters of the calling pipeline except for cov_Q of 30, var_n of 2, var_z (minimum normal coverage) of 25, var_a (minimum AS-XS) of 10, var_v (maximum normal VAF of 0.1) and indel_v (maximum normal VAF) of 0.1. After variant calling, we further filtered variants to those below 1% VAF. The few variants observed between 1% and 10% duplex VAF as variants were highly enriched for mapping artefacts, particularly for indels. No excluded variant from the additional 1% VAF threshold was found to be either a likely driver or a ClinVar pathogenic variant.<\/p>\n<p>A set of common germline variants from dbsnp<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 67\" title=\"Sherry, S. T. et al. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 29, 308&#x2013;311 (2001).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR67\" id=\"ref-link-section-d7063975e1939\" rel=\"nofollow noopener\" target=\"_blank\">67<\/a> and a custom set of known artefactual call sites in NanoSeq datasets were masked for coverage and variant calls as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e1943\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>.<\/p>\n<p>Assessing DNA contamination<\/p>\n<p>The single-molecule accuracy of the duplex sequencing method NanoSeq allows sequencing of polyclonal cell types such as sperm, but also renders mutation calls sensitive to nontarget cell-type contamination and to contamination of foreign DNA. Nontarget cell-type contamination was evaluated using manual cell counting of semen samples, which resulted in the exclusion of six samples with a sperm count of &lt;1\u2009million sperm per ml. Sperm counting methods and analysis are detailed in Supplementary Note\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>.<\/p>\n<p>Foreign DNA contamination in whole-genome NanoSeq samples was assessed using verifyBamID<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 68\" title=\"Zhang, F. et al. Ancestry-agnostic estimation of DNA sample contamination from sequence reads. Genome Res. 30, 185&#x2013;194 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR68\" id=\"ref-link-section-d7063975e1961\" rel=\"nofollow noopener\" target=\"_blank\">68<\/a>, which checks whether reads in a BAM file match previous genotypes for a specific sample, with higher values indicating more contamination. Three blood whole-genome NanoSeq samples were excluded on the basis of a verifyBamID alpha value above the suggested<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e1965\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> cut-off value of 0.005. In sperm, we found that several samples had outlier mutation burdens with verifyBamID values just below the 0.005 cut-off. This result is logical, as sperm has a much lower mutation rate compared with somatic tissues, for which the recommended cut-off was designed. Consequently, sperm samples are more sensitive to low levels of contamination. To account for this, we adjusted the verifyBamID alpha threshold for sperm to a more stringent level of 0.002, which resulted in the exclusion of three samples on the basis of this criterion.<\/p>\n<p>When assessing foreign DNA contamination in targeted and exome samples, we found that 9 targeted and 3 exome samples had verifyBamID values above &gt;0.002, slight outlier mutation burdens and a high ratio of SNP masked variants to passed variants (4-fold to 16-fold more masked variants). After further investigation, we found that all samples that exceeded verifyBamID thresholds were processed in the same sequencing batch and that this contamination could be explained by inherited germline variants of other samples in that same batch. This result suggested that a small amount of cross-contamination may have occurred during sample preparation or sequencing steps. To remove contaminant germline mutations, we performed an in silico decontamination step as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e1972\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>. This step involved calling germline variants from all targeted and exome samples using bcftools mpileup<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 69\" title=\"Danecek, P. et al. Twelve years of SAMtools and BCFtools. GigaScience 10, giab008 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR69\" id=\"ref-link-section-d7063975e1976\" rel=\"nofollow noopener\" target=\"_blank\">69<\/a> at sites where there were &gt;10 reads and a mutation call with VAF\u2009&gt;\u20090.3. All such sites were subsequently masked across all samples for both mutation calls and coverage, which essentially extended the default common SNP mask to also include rare inherited variants across the cohort. This approach resulted in all samples previously identified as contaminated with mutation burdens consistent with their age and all with a ratio of masked to passed variants of &lt;0.1, and were therefore retained for analysis.<\/p>\n<p>Corrected mutation burdens<\/p>\n<p>Given that mutation rates are strongly influenced by trinucleotide composition, it is important to consider differences in sequence composition when comparing mutation rates in datasets that target different regions of the genome<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 70\" title=\"Subramanian, S. &amp; Kumar, S. Neutral substitutions occur at a faster rate in exons than in noncoding DNA in primate genomes. Genome Res. 13, 838&#x2013;844 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR70\" id=\"ref-link-section-d7063975e1988\" rel=\"nofollow noopener\" target=\"_blank\">70<\/a>. The coding region target panels and the restriction enzyme used in whole-genome NanoSeq for instance, each will systematically bias sequencing coverage to specific genomic regions that may not reflect the full genome. To correct for these effects, in each sperm NanoSeq dataset, we generated a corrected mutation burden relative to the full genome trinucleotide frequencies as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e1992\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>.<\/p>\n<p>Comparison of burden estimates<\/p>\n<p>To compare whole-genome NanoSeq mutation burdens to mutation burdens from standard WGS, we multiplied the corrected mutation burden estimates described in the previous section by the genome size per cell type. We assumed 2,861,326,455 mappable base pairs in a haploid cell for germline datasets and the diploid equivalent for blood.<\/p>\n<p>External datasets for comparison to NanoSeq results were processed to achieve comparable burden estimates. For testis WGS samples<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 71\" title=\"Moore, J. E. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699&#x2013;710 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR71\" id=\"ref-link-section-d7063975e2007\" rel=\"nofollow noopener\" target=\"_blank\">71<\/a>, we implemented a previously described method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 4\" title=\"Abascal, F. et al. Somatic mutation landscapes at single-molecule resolution. Nature 593, 405&#x2013;410 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR4\" id=\"ref-link-section-d7063975e2011\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> that restricts analysis to regions with high coverage (20+ reads) that overlap with NanoSeq covered regions and corrected for differences in trinucleotide background frequencies relative to the full genome. For trio paternally phased DNMs from standard sequencing, as a callable genome size per sample following thorough filtering was available, we generated the mutation per cell estimate by multiplying the paternally phased DNM count by the ratio of the total genome size to the callable genome size of the sample as follows: DNMs paternal\u2009\u00d7\u2009(total genome\/callable genome). For comparison with cell types in blood, we directly compared results to published mutation burden regressions<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Machado, H. E. et al. Diverse mutational landscapes in human lymphocytes. Nature 608, 724&#x2013;732 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR13\" id=\"ref-link-section-d7063975e2015\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>.<\/p>\n<p>Mutation burden regressions<\/p>\n<p>To investigate the relationship between age and mutation burdens, we performed linear mixed-effects regression analyses. For each tissue and mutation type for which a regression was performed, the model was constructed using the lmer function from the lme4 package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Bates, D., M&#xE4;chler, M., Bolker, B. &amp; Walker, S. Fitting linear mixed-effects models using lme4. J. Stat. Softw. 67, 1&#x2013;48 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR72\" id=\"ref-link-section-d7063975e2028\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a> in R. Each model included age at sampling as a fixed effect and a random slope for each individual to account for multiple time point samples, with\u00a0restricted maximum likelihood (REML) set to false,\u00a0specified as follows:<\/p>\n<p>$${\\rm{l}}{\\rm{m}}{\\rm{e}}{\\rm{r}}({\\rm{b}}{\\rm{u}}{\\rm{r}}{\\rm{d}}{\\rm{e}}{\\rm{n}}\\sim {\\rm{a}}{\\rm{g}}{\\rm{e}}+({\\rm{a}}{\\rm{g}}{\\rm{e}}-1|{\\rm{i}}{\\rm{n}}{\\rm{d}}{\\rm{i}}{\\rm{v}}{\\rm{i}}{\\rm{d}}{\\rm{u}}{\\rm{a}}{\\rm{l}}),\\,{\\rm{R}}{\\rm{E}}{\\rm{M}}{\\rm{L}}={\\rm{F}}{\\rm{A}}{\\rm{L}}{\\rm{S}}{\\rm{E}})$$<\/p>\n<p>The 95% CIs for regression lines were calculated through bootstrapping by simulating prediction intervals. For each model, we generated 1,000 bootstrap samples. Predictions and their associated standard errors were calculated for a sequence of ages from 14 to 84 years. The 95% CIs were then derived by determining the range within which 95% of the bootstrap sample predictions fell.<\/p>\n<p>Mutational signature analysis<\/p>\n<p>We extracted DNM signatures using hierarchical dirichlet process (HDP; <a href=\"https:\/\/github.com\/nicolaroberts\/hdp\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/nicolaroberts\/hdp<\/a>), which is based on the Bayesian hierarchical dirichlet process. HDP was run with double hierarchy: (1) individual identifier (ID) and (2) tissue types (either blood or sperm), without the COSMIC reference signatures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 73\" title=\"Alexandrov, L. B. et al. Clock-like mutational processes in human somatic cells. Nat. Genet. 47, 1402&#x2013;1407 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR73\" id=\"ref-link-section-d7063975e2251\" rel=\"nofollow noopener\" target=\"_blank\">73<\/a> (v.3.3) as priors, on the mutation matrices. The number of mutations were normalized for the trinucleotide context abundance specific for each sample relative to the full genome. Both clustering hyperparameters, beta and alpha, were set to one. The Gibbs samples ran for 30,000 burn-in iterations (parameter \u2018burnin\u2019), with a spacing of 200 iterations (parameter \u2018space\u2019), from which 100 iterations were collected (parameter \u2018n\u2019). After each Gibbs iteration, three iterations of concentration parameters were conducted (parameter \u2018cpiter\u2019). Two components were extracted as DNM signatures, which were further reconstructed and decomposed into known COSMIC (v.3.3) SBS signatures using SigProfilerAssignment (<a href=\"https:\/\/github.com\/AlexandrovLab\/SigProfilerAssignment\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/AlexandrovLab\/SigProfilerAssignment<\/a>).<\/p>\n<p>Quantifying selection with dN\/dS<\/p>\n<p>To examine genes under positive selection and to quantify global selection, we used the dNdScv algorithm<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Martincorena, I. et al. Universal patterns of selection in cancer and somatic tissues. Cell 171, 1029&#x2013;1041 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR28\" id=\"ref-link-section-d7063975e2270\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>. This algorithm was extended using the base-pair-level duplex coverage, the methylation level and the pentanucleotide context to capture more complex context-dependent mutational biases and to achieve higher accuracy for our selection analysis. Detailed methods for input mutations, model selection and evaluation, site dN\/dS tests, driver mutation estimation and gene set enrichment are described in Supplementary Note\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>.<\/p>\n<p>Gene disease and mechanism annotation<\/p>\n<p>Positively selected genes were annotated with monoallelic disease consequences using the 29 February 2024 release of the DDG2P database<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 39\" title=\"Thormann, A. et al. Flexible and scalable diagnostic filtering of genomic variants using G2P with Ensembl VEP. Nat. Commun. 10, 2373 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR39\" id=\"ref-link-section-d7063975e2285\" rel=\"nofollow noopener\" target=\"_blank\">39<\/a> and the 21 June 2024 release of the OMIM database<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 74\" title=\"Amberger, J. S., Bocchini, C. A., Scott, A. F. &amp; Hamosh, A. OMIM.org: leveraging knowledge across phenotype&#x2013;gene relationships. Nucleic Acids Res. 47, D1038&#x2013;D1043 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR74\" id=\"ref-link-section-d7063975e2289\" rel=\"nofollow noopener\" target=\"_blank\">74<\/a>. OMIM annotations related to somatic disease, complex disease or tentative disease associations were excluded.<\/p>\n<p>Genes were also annotated for their potential mutation mechanism observed in sperm and cancer or developmental disorders when available. In sperm, genes were labelled as LOF if they had nominal enrichment of nonsense\u2009+\u2009splice variants and\/or indel variants (ptrunc_cv &lt;0.1 | pind_cv &lt;0.1) and 2+ LOF mutations. There were two exceptions to this, whereby genes met these thresholds but were labelled as activating owing to having a restricted repertoire of LOF mutations that are known to be oncogenic in cancers: CBL (LOFs in and downstream of the RING zinc finger domain)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 75\" title=\"Martinelli, S. et al. Molecular diversity and associated phenotypic spectrum of germline CBL mutations. Hum. Mutat. 36, 787&#x2013;796 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR75\" id=\"ref-link-section-d7063975e2299\" rel=\"nofollow noopener\" target=\"_blank\">75<\/a> and PPM1D (LOFs in final two exons)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 76\" title=\"Khadka, P. et al. PPM1D mutations are oncogenic drivers of de novo diffuse midline glioma formation. Nat. Commun. 13, 604 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR76\" id=\"ref-link-section-d7063975e2306\" rel=\"nofollow noopener\" target=\"_blank\">76<\/a>. All other genes had missense enrichment only and were labelled as activating. The mechanism in cancer was defined by using the \u2018role in cancer\u2019 field of the COSMIC cancer gene census (v.99)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 38\" title=\"Sondka, Z. et al. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat. Rev. Cancer 18, 696&#x2013;705 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR38\" id=\"ref-link-section-d7063975e2310\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a>, where \u2018oncogene\u2019 was labelled as activating and \u2018tumour suppressor gene\u2019 as LOF. Annotations of a fusion mechanism were not displayed, except for genes that had neither an oncogene nor a tumour suppressor annotation, which were labelled as \u2018fusion only\u2019. The developmental disorder mechanism was defined by using the variation consequence field of DDG2P, for which \u2018restricted repertoire of mutations;activating\u2019 was labelled as activating and \u2018loss_of_function_variant\u2019 was labelled as LOF.<\/p>\n<p>Gene mutation plots<\/p>\n<p>The lollipop gene mutation plots were created using a coordinate system, whereby position\u20091 was defined as the first coding base of the GRCh37-GencodeV18+Appris<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 77\" title=\"Frankish, A. et al. GENCODE 2021. Nucleic Acids Res. 49, D916&#x2013;D923 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR77\" id=\"ref-link-section-d7063975e2322\" rel=\"nofollow noopener\" target=\"_blank\">77<\/a> transcript of the gene. The data sources included protein domains from UniProt<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"The UniProt Consortium. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523&#x2013;D531 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR78\" id=\"ref-link-section-d7063975e2326\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a>, somatic mutations from the exome and genome-wide screens of the COSMIC (v.99)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 38\" title=\"Sondka, Z. et al. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat. Rev. Cancer 18, 696&#x2013;705 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR38\" id=\"ref-link-section-d7063975e2330\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a>, ClinVar (release 2024.07.01)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062&#x2013;D1067 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR41\" id=\"ref-link-section-d7063975e2334\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a> pathogenic annotation, per base pair cohort-wide (targeted\u2009+\u2009exome) NanoSeq coverage, sperm mutation count (number of independent individuals with a mutation) and mutation consequence and amino acid change annotated by the dNdScv algorithm<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Martincorena, I. et al. Universal patterns of selection in cancer and somatic tissues. Cell 171, 1029&#x2013;1041 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR28\" id=\"ref-link-section-d7063975e2338\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>. These data were plotted using code modified from the lolliplot function in the trackViewer R\u2009package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"Ou, J. &amp; Zhu, L. J. trackViewer: a Bioconductor package for interactive and integrative visualization of multi-omics data. Nat. Methods 16, 453&#x2013;454 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR79\" id=\"ref-link-section-d7063975e2343\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>.<\/p>\n<p>Variant annotation<\/p>\n<p>Variants were annotated using Ensembl\u2019s variant effect predictor (VEP)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"McLaren, W. et al. The Ensembl Variant Effect Predictor. Genome Biol. 17, 122 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR80\" id=\"ref-link-section-d7063975e2355\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a> with added custom annotations of mutation context, ClinVar (release 2024.07.01)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062&#x2013;D1067 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR41\" id=\"ref-link-section-d7063975e2359\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a>, CADD (v.GRCh37-v1.6)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Rentzsch, P., Schubach, M., Shendure, J. &amp; Kircher, M. CADD-Splice&#x2014;improving genome-wide variant effect prediction using deep learning-derived splice scores. Genome Med. 13, 31 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR42\" id=\"ref-link-section-d7063975e2363\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a> and the average methylation level in testis. Methylation data were obtained from whole-genome shotgun bisulfite sequencing methylation data from male testis samples from a 37-year-old (ENCFF638QVP) and 54-year-old (ENCFF715DMX) from the ENCODE project<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 71\" title=\"Moore, J. E. et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 583, 699&#x2013;710 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR71\" id=\"ref-link-section-d7063975e2367\" rel=\"nofollow noopener\" target=\"_blank\">71<\/a>. The average methylation level was calculated by selecting CpG sites with coverage of three or more and averaging the per\u2009cent of sites methylated between the two samples.<\/p>\n<p>Variants were annotated as likely monoallelic disease-causing mutations if they met at least one of two criteria: (1) reported in ClinVar as pathogenic, likely pathogenic or if they were reported as \u2018conflicting classifications of pathogenicity\u2019, where the conflict was between reports of pathogenic\/likely pathogenic and \u2018uncertain significance\u2019 with no reports of benign or likely benign and not specified as a recessive condition; or (2) were a highly damaging variant in high-confidence monoallelic developmental disorder genes from DDG2P<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 39\" title=\"Thormann, A. et al. Flexible and scalable diagnostic filtering of genomic variants using G2P with Ensembl VEP. Nat. Commun. 10, 2373 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR39\" id=\"ref-link-section-d7063975e2374\" rel=\"nofollow noopener\" target=\"_blank\">39<\/a>. Genes met the following criteria: (1) allelic requirement being monoallelic_autosomal, monoallelic_X_hem, monoallelic_X_het or mitochondrial; (2) confidence in strong, definitive or moderate; and (3) a mutation consequence of \u2018absent gene product\u2019. Highly damaging was defined as being a \u2018high\u2019 impact variant in the VEP annotation (frameshift splice_acceptor, splice_donor, start_lost, stop_gained or stop_lost) or a missense variant with CADD<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Rentzsch, P., Schubach, M., Shendure, J. &amp; Kircher, M. CADD-Splice&#x2014;improving genome-wide variant effect prediction using deep learning-derived splice scores. Genome Med. 13, 31 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#ref-CR42\" id=\"ref-link-section-d7063975e2378\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a> score &gt;30 (top 0.1% damaging).<\/p>\n<p>Variants were defined as a likely driver if they met the \u2018highly damaging\u2019 criteria defined above in a significant germline-selection gene with LOF mutation enrichment or if they were in 1 of the 24 significant mutation hotspots. This resulted in 320 variants being labelled as likely drivers in exome samples.<\/p>\n<p>Cell fraction mutation estimates<\/p>\n<p>To calculate the mean count of synonymous, missense or LOF or pathogenic mutations per sperm cell, we summed the duplex VAFs of all variants in that class. For example, if an individual had 3 synonymous mutations each observed once with a duplex coverage of 100 at each of those sites, each of those variants would have a duplex VAF of 1\/100\u2009=\u20090.01. The sum of VAFs in this example would then be 0.03 and this would then be reported as the estimate for the mean count of synonymous variants per sperm cell for that individual. At low fractions such as 0.03, the mean count per cell is approximately equivalent to the percentage of sperm with this mutation class (3%); therefore, the driver and disease mutations are reported as percentage estimates. At higher fractions (for example, mean count\u2009&gt;\u20091), the fractions are not equivalent to percentage, as many cells will have multiple variants of that class; therefore, the estimates are reported as mean count.<\/p>\n<p>Expected mean counts for SNVs were generated by annotating each possible substitution at each covered site with an expected number of mutations per sample as given by expCountSNV\u2009=\u2009context_mut_rate\u2009\u00d7\u2009duplex_coverage\u2009\u00d7\u2009age_correction. The context_mut_rate was given by the 208 basePairCov\u2009+\u2009cpgMeth trinucleotide mutation model estimates for that trinucleotide\u2009+\u2009methylation mutation context (Supplementary Note\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM3\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>). Duplex coverage is the exact duplex coverage at that site for that sample. The age_correction parameter was given to normalize the mutation model estimates (derived from all exome samples) to the mutation rate of that sample based on age. Specifically, we fit a linear model to the mutation burden versus age of the exome samples and used this to generate a predicted mutation rate for each sample based on age. The per sample correction parameter was calculated as the age predicted mutation burden divided by the mean mutation rate of all exome samples (3.89\u2009\u00d7\u200910\u20138). The resulting corrections spanned from 0.50 (youngest sample) to 1.42 (oldest samples). The expected indel mutation rate was calculated in the same way, except with a single mutation rate parameter (indels\/bp) expCountIndel\u2009=\u2009indel_mut_rate\u2009\u00d7\u2009duplex_coverage\u2009\u00d7\u2009age_correction. The expected mean count was then calculated for each category (for example, synonymous or likely disease) by summing the expected values for each SNV and\/or indel base pair matching the relevant annotation. As background for possible ClinVar pathogenic variants, we only considered indels of size 21\u2009bp or less, the largest detected indel in the dataset. Regressions were fit with either linear models or generalized linear models (glm in R) with family\u2009=\u2009quasibinomial.<\/p>\n<p>Regression analysis<\/p>\n<p>We tested for associations between mutation outcome variables from sperm genome, sperm exome, sperm targeted and blood genome NanoSeq data and the phenotype predictor variables of BMI, smoking pack-years and alcohol drink-years. These tests were performed using a Gaussian family generalized linear regression in R. For each mutation outcome variable the test took the following form: glm(mutationOutcome\u2009~\u2009age_at_sampling\u2009+\u2009BMI\u2009+\u2009pack_years\u2009+\u2009drinkYears, family\u2009=\u2009\u201cgaussian\u201d).<\/p>\n<p>The mutation outcome variables examined were SNV and indel burden from all four sequencing datasets, SBS1 and SBS5 count from sperm genomes, SBS1, SBS5 and SBS19 from blood genomes, and likely disease cell fraction and likely driver cell fraction from sperm targeted and sperm exomes. The significance of each predictor was assessed from the summary output coefficients of the model, and P\u2009values were adjusted for 68 total tests using the false discovery rate method.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-025-09448-3#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Ethics This study was carried out under TwinsUK BioBank ethics, approved by the North West\u2013Liverpool Central Research Ethics&hellip;\n","protected":false},"author":2,"featured_media":217142,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25],"tags":[49,48,104878,316,37832,1099,1100,44874,66],"class_list":["post-217141","post","type-post","status-publish","format-standard","has-post-thumbnail","category-genetics","tag-ca","tag-canada","tag-disease-genetics","tag-genetics","tag-genome-evolution","tag-humanities-and-social-sciences","tag-multidisciplinary","tag-rare-variants","tag-science"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/217141","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=217141"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/217141\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/217142"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=217141"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=217141"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=217141"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}