{"id":225519,"date":"2025-10-19T22:34:20","date_gmt":"2025-10-19T22:34:20","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/225519\/"},"modified":"2025-10-19T22:34:20","modified_gmt":"2025-10-19T22:34:20","slug":"genotyping-sequence-resolved-copy-number-variation-using-pangenomes-reveals-paralog-specific-global-diversity-and-expression-divergence-of-duplicated-genes","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/225519\/","title":{"rendered":"Genotyping sequence-resolved copy number variation using pangenomes reveals paralog-specific global diversity and expression divergence of duplicated genes"},"content":{"rendered":"<p>Overview of the genotyping method<\/p>\n<p>We represent variation as haplotype segments that are short enough to minimize disruption by recombination, allowing precise sharing with an NGS sample through identity by descent<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Browning, S. R. &amp; Browning, B. L. Identity by descent between distant relatives: detection and applications. Annu. Rev. Genet. 46, 617&#x2013;633 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR23\" id=\"ref-link-section-d26018869e476\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a> while capturing structural information including phased small variation, structural variation and gene conversion events. We label these haplotype segments as pangenome-derived alleles (PAs) and detect PAs shared with an NGS sample. PA boundaries are arranged to study variation of protein-coding genes. Each PA includes consecutive exons separated by &lt;20\u2009kb and 5\u2009kb of flanking sequences (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), reflecting functional proximity of short-range transcription factors<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 24\" title=\"Chen, C.-H. et al. Determinants of transcription factor regulatory range. Nat. Commun. 11, 2472 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR24\" id=\"ref-link-section-d26018869e483\" rel=\"nofollow noopener\" target=\"_blank\">24<\/a> and population-level genomic linkage. PAs typically range between 10 and 100\u2009kb, corresponding to the scale of linkage disequilibrium (LD) blocks<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 25\" title=\"Wall, J. D. &amp; Pritchard, J. K. Haplotype blocks and linkage disequilibrium in the human genome. Nat. Rev. Genet. 4, 587&#x2013;597 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR25\" id=\"ref-link-section-d26018869e487\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>. While PAs generally correspond to individual genes, they also cover fractions of genes with long introns or, conversely, include tandemly arrayed paralogs within 20\u2009kb.<\/p>\n<p>For computational efficiency and to avoid alignment ambiguity in repetitive DNA, we use an alignment-free comparison of low-copy k-mers (DNA fragments of a fixed length k; k\u2009=\u200931) measured in NGS samples to genotype PAs. For each gene, we group all similar PAs in the pangenome, including orthologs, paralogs and homologous pseudogenes, and construct a matrix used in genotyping that contains the k-mer composition of all grouped PAs (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). The rows of a matrix correspond to individual PAs, columns correspond to k-mers exclusive to the grouped PAs, and cell values represent the k-mer multiplicity in each PA (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1a,b<\/a>).<\/p>\n<p>Fig. 1: Overview of the genotyping method.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41588-025-02346-4\/figures\/1\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig1\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2025\/10\/41588_2025_2346_Fig1_HTML.png\" alt=\"figure 1\" loading=\"lazy\" width=\"685\" height=\"324\"\/><\/a><\/p>\n<p>a, Demography of the reference pangenome assemblies. Map credit: Hogweard\/Wikimedia Commons. b, Construction of pangenome k-mer matrices for CNV genes. Each individual gene is represented as a vector of counts of k-mers exclusively found among homologous sequences. All similar sequences including paralogs and orthologs are included and integrated as a k-mer matrix. c, Construction of phylogenetic trees based on k-mer matrices. d, Schematic of approach to estimate genotypes of alleles using NGS data. The k-mers from each matrix are counted in NGS data and normalized by sequencing depth. The normalized k-mer counts are projected to all pangenome genes. e, Reprojection of the raw results in the last step to integer solutions recursively based on the phylogenetic tree. f, An illustrative annotation and genotyping results on the SMN1 and SMN2 genes using HPRC samples. On the left side of the classification, the phylogenetic tree and heatmap of pairwise similarities are shown along with a mutation plot based on an MSA highlighting point differences in SMN1 in CHM13. All SMN genes are categorized into five major types and 17 subgroups. SMN1 and SMN2 correspond to the most common types of each paralog; SMN1-2, SMN1 partially converted to SMN2; SMN-conv, converted SMN genes, predominantly mapped to the SMN2 locus and found to be enriched in African populations; SMN2-2, a rare outgroup of SMN2. The GRCh38 assembly includes SMN1-2 and SMN2. The Phe280 T variant that disrupts the splicing of SMN2 transcripts is highlighted in red. The genotyping results of 1kGP continental populations are shown on the right. Rows correspond to subgroups, columns correspond to continental populations, and the colors of pie charts give distributions of copy numbers (CNs) among each continental population. AFR, African; AMR, American; EAS, East Asian; EUR, European; SAS, South Asian; ref, reference.<\/p>\n<p>The genotyping is performed per matrix by identifying a combination of PAs (rows) and their copy number with the least-squared distance between their k-mer counts and that from an NGS sample. The sample k-mer counts are projected into the vector space of each k-mer matrix and assigned integer copy numbers using recursive rounding based on the phylogeny of PA sequences (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1c\u2013e<\/a>), resulting in a list of PA-specific copy numbers (paCNs). For example, there are 178 PAs for SMN genes, the gene family associated with spinal muscular atrophy. This includes copies of SMN1, SMN2 and paralogs with gene conversion<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 26\" title=\"Ogino, S., Gao, S., Leonard, D. G. B., Paessler, M. &amp; Wilson, R. B. Inverse correlation between SMN1 and SMN2 copy numbers: evidence for gene conversion from SMN2 to SMN1. Eur. J. Hum. Genet. 11, 275&#x2013;277 (2003).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR26\" id=\"ref-link-section-d26018869e648\" rel=\"nofollow noopener\" target=\"_blank\">26<\/a>, for example, paralogs mapped to SMN2 that contain the SMN1 version of Phe280, the single nucleotide polymorphism (SNP) responsible for dysfunctional exon 7 splicing of SMN2 (ref.\u2009<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Lorson, C. L., Hahnen, E., Androphy, E. J. &amp; Wirth, B. A single nucleotide in the SMN gene regulates splicing and is responsible for spinal muscular atrophy. Proc. Natl Acad. Sci. USA96, 6307&#x2013;6311 (1999).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR27\" id=\"ref-link-section-d26018869e662\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>) (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1f<\/a>).<\/p>\n<p>PA database construction<\/p>\n<p>We constructed a PA database for 3,351 genes previously reported as CNVs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312&#x2013;324 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR18\" id=\"ref-link-section-d26018869e677\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Gao, Y. et al. A pangenome reference of 36 Chinese populations. Nature 619, 112&#x2013;121 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR20\" id=\"ref-link-section-d26018869e680\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a> (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>), using 114 diploid PacBio HiFi assemblies from the Human Pangenome Reference Consortium (HPRC), the Human Genome Structural Variation Consortium (HGSVC), the Chinese Pangenome Consortium (CPC) and two telomere-to-telomere assemblies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"He, Y. et al. T2T-YAO: a telomere-to-telomere assembled diploid reference genome for Han Chinese. Genom. Proteom. Bioinform. 21, 1085&#x2013;1100 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR28\" id=\"ref-link-section-d26018869e687\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Yang, C. et al. The complete and fully-phased diploid genome of a male Han Chinese. Cell Res. 33, 745&#x2013;761 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR29\" id=\"ref-link-section-d26018869e690\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>, in addition to GRCh38 and CHM13 (ref.\u2009<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Nurk, S. et al. The complete sequence of a human genome. Science 376, 44&#x2013;53 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR30\" id=\"ref-link-section-d26018869e694\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>). In total, we defined 1,408,209 PAs, organized into 3,307 matrices (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a\u2013c<\/a>).<\/p>\n<p>Fig. 2: Overview of the database of PAs.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41588-025-02346-4\/figures\/2\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig2\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2025\/10\/41588_2025_2346_Fig2_HTML.png\" alt=\"figure 2\" loading=\"lazy\" width=\"685\" height=\"849\"\/><\/a><\/p>\n<p>a, An example of amylase 1 PAs. Left: the corresponding order of all AMY1 PAs on assemblies, which are colored based on their major groups (unions of subgroups devoid of SVs &gt;300\u2009bp between members). Right: AMY1 genes are extracted as PAs as well as their flanking genes and sequences, including AMY2B translocated proximally to AMY1 and two pseudogenes: AMYP1 and RP5-1108M17. All PAs are vertically ordered according to the phylogenetic tree and aligned via graphic MSAs (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>). Homologous sequences are vertically aligned. Mutations are visualized as dots, and large gaps (deletions) are shown as spaces. Seven major groups are categorized including five paralogs and two orthologs. There are no pseudogenes around AMY1C, while AMY1A has RP5-1108M17 nearby and AMY1B has AMYP1 nearby. There are alternative versions of AMY1B and AMY1C with sequence substitutions. A new paralog called AMY1(Dup) is found primarily on haplotypes with duplications and has both pseudogenes nearby. Another paralog of AMY1 found with translocated AMY2B is called AMY1\u2009+\u2009AMY2B. There are also two rare paralogs (blue and violet) and one singleton ortholog (steel blue). Alt, alternative; dup, Duplication. b, The size distribution of PAs on a log density. c, Circos plot of all PAs. Outer ring, the density of PAs in each Mb on GRCh38. Arcs, interchromosomal PAs are included in the same group. d, Annotation of PAs according to orthology and variants with respect to GRCh38. e, Identifiability of highly similar subgroups by unique k-mers. The total number of subgroups (blue) and the number of subgroups that may be identified by paralog-specific k-mers (red) are shown for each matrix with a size of at least three. f, Distribution of logistic pairwise divergence of PAs depending on orthology and phylogenetic relationship. The values shown are average values from each matrix. Small neighbor distances are an indicator of the strong representativeness of the current cohort. g, Saturation analysis for all subgroups using a recapture mode according to two sorted orders: African genomes considered first and non-African genomes considered first. The former order has a smoother curve than the latter, indicating that there are more African-specific subgroups.<\/p>\n<p>Because of limited human genetic diversity and stronger LD across short distances, PAs are often highly similar or identical. To reduce dimensionality and facilitate cohort analysis, we used their phylogenetic relationships to merge similar PAs into highly similar subgroups (subgroups) treated as equal states (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). In total, we defined 89,236 subgroups, which were used to enumerate all PAs, analogous to human leukocyte antigen (HLA) nomenclature (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>).<\/p>\n<p>To annotate low-frequency variants and reference genome locations for orthologous or paralogous relationships, we mapped PAs to GRCh38 (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>). In total, 164,237 paralogous PAs across 6,389 loci were determined. Paralogous PAs that were similar to their corresponding reference locus (\u226580% k-mer similarity) were labeled duplicative, and the remaining lower-identity paralogous PAs were labeled diverged. In total, 10,792 diverged paralogs from 2,734 subgroups were identified across 333 matrices (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2d<\/a>). The divergent paralogs represent new sequences recalcitrant to canonical reference analysis. For example, some amylase PAs include paralogs for both AMY1 and AMY2B, reflecting an AMY2B translocation (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>).<\/p>\n<p>While most duplications were distal to their original genes, 6,673 PAs reflected proximal (&lt;20\u2009kb) duplications, including 1,646 PAs across 36 genes exhibiting \u2018runaway duplication\u2019 (ref.\u2009<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Handsaker, R. E. et al. Large multiallelic copy number variations in humans. Nat. Genet. 47, 296&#x2013;303 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR31\" id=\"ref-link-section-d26018869e837\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>) with at least three proximal duplications (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). Proximally duplicated genes were included in the same PA as their ortholog as a heritable unit. Orthologous PAs were classified as reference alleles if they belonged to the same subgroup as the reference gene and as alternative alleles otherwise (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2d<\/a>). All PAs were genotyped regardless of paralog\u2013ortholog annotation so that the resulting genotypes contain population and copy number variation.<\/p>\n<p>Ctyper databases capture population diversity<\/p>\n<p>We assessed whether PAs capture unique aspects of genomic information that cannot be replicated by other CNV representations, including copy numbers of reference genes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Sudmant, P. H. et al. Diversity of human copy number variation and multicopy genes. Science 330, 641&#x2013;646 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR1\" id=\"ref-link-section-d26018869e855\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Handsaker, R. E. et al. Large multiallelic copy number variations in humans. Nat. Genet. 47, 296&#x2013;303 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR31\" id=\"ref-link-section-d26018869e858\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>, singly unique nucleotide k-mers<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 1\" title=\"Sudmant, P. H. et al. Diversity of human copy number variation and multicopy genes. Science 330, 641&#x2013;646 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR1\" id=\"ref-link-section-d26018869e865\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a> (SUNKs) and large haplotype structures<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Yilmaz, F. et al. Reconstruction of the human amylase locus reveals ancient duplications seeding modern-day variation. Science 386, eadn0609 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR32\" id=\"ref-link-section-d26018869e869\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Bolognini, D. et al. Recurrent evolution and selection shape structural diversity at the amylase locus. Nature 634, 617&#x2013;625 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR33\" id=\"ref-link-section-d26018869e872\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>. We found that PAs provide higher resolution of variation (for example, single-nucleotide variants), as 94.7% of variants are not reflected by sequences in GRCh38. Additionally, both nearby SUNK markers (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2e<\/a>) and large haplotype structures were found to be poor proxies for PAs, and only a small proportion of PAs were found to link to SUNKs or larger haplotypes (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). Despite largely reduced dimensions, subgroups capture more than 80% of the total population variation (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2f<\/a>). Finally, using saturation analysis<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Wong, K. H. Y. et al. Towards a reference genome that captures global genetic diversity. Nat. Commun. 11, 5482 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR34\" id=\"ref-link-section-d26018869e889\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a>, we estimate that the current cohort represents 98.7% of subgroups in non-Africans and 94.9% in Africans (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), suggesting a near-saturated database (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2g<\/a>).<\/p>\n<p>Benchmarking genotypes from NGS samples<\/p>\n<p>We genotyped 2,504 unrelated individuals and 641 offspring from the 1000 Genomes Project (1kGP). Most subgroups (99.25%) showed Hardy\u2013Weinberg equilibrium (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig.\u2009<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3a<\/a>) and thus little bias. There were 27 matrices with &gt;15% subgroups in disequilibrium, which were mostly short genes (median\u2009=\u20094,564\u2009bp) with few low-copy k-mers (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). Genotypes were accurate with an average F1 score for trio concordance of 97.58% (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>, Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3b<\/a>), while 18 matrices had high discordance (&gt;15%), primarily for subtelomeric genes or on sex chromosomes with poorer assembly qualities (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>).<\/p>\n<p>Fig. 3: Benchmarking of genotyping results.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41588-025-02346-4\/figures\/3\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig3\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2025\/10\/41588_2025_2346_Fig3_HTML.png\" alt=\"figure 3\" loading=\"lazy\" width=\"685\" height=\"762\"\/><\/a><\/p>\n<p>a, Hardy\u2013Weinberg equilibrium of genotyping results from 2,504 1kGP unrelated samples. b, Genotype concordance of genotyping results from 641 1kGP trios, ordered by F1 error. The matrices with an F1 error of more than 15% were labeled by genomic location. c, Copy number comparison between assemblies and genotyping results from 39 HPRC samples shared with 1kGP. d, Sequence differences between genotyped and original alleles during the leave-one-out test using pairwise alignment of nonrepetitive sequences. e, Detailed leave-one-out comparison in the diploid telomere-to-telomere genome CN1. The results are categorized regarding the number of paralogs in CHM13 to show performances on different levels of genome complexity and the main sources of errors. The CN1 NGS sample had about 40-fold coverage. f, Ctyper runtime on all loci in CN1 for varying coverage. g, Benchmarking of HLA genotyping using ctyper on full and leave-one-out (LOO) databases, compared with T1K on 31 HLA genes. h, Benchmarking of CYP2D annotation on all CYP2D genes and CYP2D6 exclusively. FN, false negative; FP, false positive; TP, true positive.<\/p>\n<p>We assessed copy number accuracy and bias among highly duplicated gene families (for example, amylase, NBPF, GOLGA and TBC1D3). The copy numbers derived from genotyping were compared to those from corresponding assemblies for 39 HPRC samples shared with the 1kGP using a database inclusive of these samples. To limit compounded error from misassembled sequences, we excluded samples with low-confidence sequences (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). For each sample, we benchmarked on all matrices for which the corresponding assembly was high in copy number (&gt;10). The copy numbers were highly correlated (\u03c1\u2009=\u20090.996, Pearson correlation) with little bias (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3c<\/a>), 0.2% missing copies (false negatives) and 2.4% additional copies (false positives), likely from unassembled genes in assemblies. High concordances remained when tests were expanded to all genotyped genes (\u03c1\u2009&gt;=\u20090.996, Pearson correlation).<\/p>\n<p>We assessed the sequence similarity of the genotyped alleles to the ground truth genome assembly for the 39 HPRC benchmarking genomes. Each sample was genotyped with the full database (full-set) or the database excluding its corresponding PAs (leave-one-out). We matched the genotyped PAs to the corresponding assembly PAs (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), excluding introns and decoys and sequences with &lt;1\u2009kb of nonrepetitive bases, and measured the similarity between the genotyped allele and the assigned query. We performed a similar analysis, treating the closest neighbor to each assembly PA from the database as the correctly genotyped locus. Due to mismatching from database sampling or misassemblies, 2.9% of PAs from the leave-one-out experiment and 1.0% from the full-set experiment were not paired with truth copies for assessment. For the full set, paired PAs had 0.36 mismatches per 10\u2009kb, with 93.0% having no mismatches in nonrepetitive regions. The leave-one-out tests had 2.7 mismatches per 10\u2009kb in nonrepetitive regions, which was 1.2 additional mismatches per 10\u2009kb from the optimal solutions (closest neighbors); 57.3% of alleles had no mismatches, and 77.0% were mapped to the optimal solution (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3d<\/a>). The leave-one-out results were 96.5% more similar to the original PAs than the closest GRCh38 gene at 79.3 mismatches per 10\u2009kb.<\/p>\n<p>To isolate sources of errors in cases of misassemblies, we directly compared leave-one-out genotyping results to a telomere-to-telomere assembly<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"He, Y. et al. T2T-YAO: a telomere-to-telomere assembled diploid reference genome for Han Chinese. Genom. Proteom. Bioinform. 21, 1085&#x2013;1100 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR28\" id=\"ref-link-section-d26018869e1032\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a> of genic PAs. The sample genotypes had 11,627 correctly matched subgroups, 599 (4.8%) mistyped to other subgroups, 131 from subgroups unique to the assembly (1.1%; out of reference), 127 false positives (0.5% F1) and 93 false negatives (0.4% F1) for a total F1 error of 6.7% (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3e<\/a>), with a copy number agreement of 99.1%. This is a 3% increase in mistypes compared to trio discordance.<\/p>\n<p>The computational requirements are sufficient for biobank analysis. The average runtime for genotyping 3,351 genes at 30\u00d7 coverage was 80.2\u2009min (1.0\u2009min per 1\u00d7 coverage for sample preprocessing and 0.9\u2009s per gene for genotyping) on a single core (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3f<\/a>) using ~20\u2009GB of RAM, with support for parallel processing.<\/p>\n<p>We compared the HLA, KIR and CYP2D6 genotypes to the locus-specific methods T1K<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 35\" title=\"Song, L., Bai, G., Liu, X. S., Li, B. &amp; Li, H. Efficient and accurate KIR and HLA genotyping with massively parallel sequencing data. Genome Res. 33, 923&#x2013;931 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR35\" id=\"ref-link-section-d26018869e1068\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a> and Aldy<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Numanagi&#x107;, I. et al. Allelic decomposition and exact genotyping of highly polymorphic and structurally variant genes. Nat. Commun. 9, 828 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR36\" id=\"ref-link-section-d26018869e1072\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>. For 31 HLA genes, ctyper had an F1 score of 98.9% across all four fields of HLA nomenclature<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Robinson, J. et al. IPD-IMGT\/HLA Database. Nucleic Acids Res. 48, D948&#x2013;D955 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR37\" id=\"ref-link-section-d26018869e1080\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 38\" title=\"Lefranc, M. P. IMGT, the international immunogenetics database. Nucleic Acids Res. 29, 207&#x2013;209 (2001).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR38\" id=\"ref-link-section-d26018869e1083\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a> against the full-set analysis and a score of 86.3% for the leave-one-out analysis, while T1K had 70.8%. For protein-coding products (first two fields), ctyper reached 99.98% against the full-set analysis (with 99.9% copy number F1 correctness) and 96.5% (with 99.5% copy number F1 correctness) for the leave-one-out analysis, and T1K had 97.2% (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3g<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>). For 14 KIR genes, ctyper reached 98.5% across all fields in the full-set analysis and 70.6% for the leave-one-out analysis, while T1K had 32.0% due to the limited database. For protein-coding products (first three digits), ctyper reached 99.2% against the full-set analysis (with 99.9% copy number F1 correctness) and 88.8% for the leave-one-out analysis (with 99.2% copy number F1 correctness), while T1K had 79.6% (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>). Benchmarking CYP2D6 star annotations of assemblies<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 39\" title=\"Pacific Biosciences. Pangu &#010;                https:\/\/github.com\/PacificBiosciences\/pangu&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR39\" id=\"ref-link-section-d26018869e1120\" rel=\"nofollow noopener\" target=\"_blank\">39<\/a>, ctyper reached 100.0% against the full-set analysis and 83.2% for the leave-one-out analysis, compared to 80.0% using Aldy (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3h<\/a>). There was perfect agreement of SNP variants for ctyper against the full-set analysis and 95.7% for the leave-one-out analysis, compared to 85.2% using Aldy.<\/p>\n<p>Finally, we used ctyper to genotype 273 CMR genes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 40\" title=\"Wagner, J. et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat. Biotechnol. 40, 672&#x2013;680 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR40\" id=\"ref-link-section-d26018869e1131\" rel=\"nofollow noopener\" target=\"_blank\">40<\/a>. Unrepetitive regions averaged 0.29 mismatches per 10\u2009kb against the full-set analysis, 99.7% fewer than when comparing assemblies to corresponding GRCh38 sequences (baseline). The genotypes using leave-one-out databases had 4.9 mismatches per 10\u2009kb, 94.8% fewer than baseline (Supplementary Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>\u2013<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>). Including repeat-masked low-complexity sequences (for example, variable-number tandem repeats), there were 10.5 mismatches per 10\u2009kb against the full-set analysis (97.6% fewer than baseline) and 74.7 mismatches per 10\u2009kb for the leave-one-out analysis (82.7% fewer than baseline; Supplementary Figs. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>\u2013<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>).<\/p>\n<p>We compared genotyping of HLA and CMR genes to a contemporary method using pangenomes, Locityper<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 41\" title=\"Prodanov, T. et al. Locityper: targeted genotyping of complex polymorphic genes. Nat. Genet. &#010;                https:\/\/doi.org\/10.1038\/s41588-025-02362-4&#010;                &#010;               (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR41\" id=\"ref-link-section-d26018869e1150\" rel=\"nofollow noopener\" target=\"_blank\">41<\/a>, using leave-one-out analysis. For HLA, Locityper achieved an F1 score of 87.9% (versus ctyper, 86.3%) for predicting all four nomenclature fields, while ctyper performed slightly better on the first two fields for protein-coding variants (96.5% versus 94.0%; <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>), although ctyper had a roughly 218\u00d7 speedup due to alignment-free genotyping. When analyzing CMR genotypes, there were 19.8 fewer mismatches per 10\u2009kb than the Locityper genotypes in comparable regions (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>, Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a> and Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>).<\/p>\n<p>Sequence-level diversity of CNVs in global populations<\/p>\n<p>We used principal-component (PC) analysis (PCA) to examine the population structure of PA genotypes in the 2,504 unrelated 1kGP samples, 879 Genotype\u2013Tissue Expression (GTEx) samples and 105 diploid assemblies (excluding HGSVC due to quality filtering), excluding rare subgroups (&lt;0.05 allele frequency) and limiting copy number to ten to balance the weights of PCs (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>). All data cluster by population as opposed to source, suggesting little bias between genotyping and assembly or across NGS cohorts. The top 0.1% highest-weighted subgroups in PC1 have an average aggreCN variance of 26.33, significantly larger than the overall of 4.00 (P value\u2009=\u20091.11\u2009\u00d7\u200910\u221216, F-test). Similarly, PC2 and PC3 have mean aggreCN variances of 19.73 and 7.20, suggesting that CNVs are weakly associated with sequence variants. Furthermore, PC1 is the only PC that clustered all samples into the same sign with a geographic center away from 0, suggesting that it corresponds to modulus variance (hence aggreCN) if treating samples as vectors of paCNs. Meanwhile, PC2 and PC3 were similar to the PCA plots based on SNP data of global samples<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Mallick, S. et al. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature 538, 201&#x2013;206 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR42\" id=\"ref-link-section-d26018869e1191\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a>, suggesting that they are associated with the sequence diversity of CNV genes. The total number of duplications is elevated in African populations (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4b<\/a>), reflected in the order of PC1 (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>).<\/p>\n<p>Fig. 4: Global population diversity in allele-specific copy number variation.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41588-025-02346-4\/figures\/4\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig4\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2025\/10\/41588_2025_2346_Fig4_HTML.png\" alt=\"figure 4\" loading=\"lazy\" width=\"685\" height=\"475\"\/><\/a><\/p>\n<p>a, PCAs of allele-specific copy numbers on the union of PA genotyping results and assembly annotations. b, Distribution of total autosomal gene copy numbers among unrelated 1kGP samples, including pseudogenes (AFR, n\u2009=\u2009685; admixed AMR, n\u2009=\u2009352; EAS, n\u2009=\u2009511; EUR, n\u2009=\u2009522; SAS, n\u2009=\u2009516; box shows median and interquartile range; whiskers extend to 1.5\u00d7 interquartile range, with outliers beyond). c, Population differentiation measured by F statistics of duplications among different continental populations. Genes with a paralogous subgroup with an F statistic of more than 0.35 are labeled. d, Mean absolute variation in copy numbers and RPD in sequences. Based on our genotyping results from unrelated 1kGP genomes, for genes found to be CNV to the population median in more than 20 samples, we determined the average aggreCN difference (MAE) between individuals and estimated the average paralog difference in sequences relative to the ortholog difference. e, mLD between pairs of CNV genes less than 100\u2009kb apart. The larger MAE value of each pair is used for the x-axis values. The total locus length denotes the length from the beginning of the first gene to the end of the last gene.<\/p>\n<p>We examined ctyper genotypes to measure the extent to which duplications show population specificity. We used the F statistic, a generalization of FST that accommodates more than two genotypes (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), to test the differences in distributions across continental populations (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4c<\/a>). In total, 4.4% (223 of 5,065) of duplicated subgroups showed population specificity (F statistic\u2009&gt;\u20090.2; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>). The subgroups of PAs with the highest F statistic (0.48) contain duplications of HERC2P9, a known differentiated gene<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Sudmant, P. H. et al. An integrated map of structural variation in 2,504 human genomes. Nature 526, 75&#x2013;81 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR7\" id=\"ref-link-section-d26018869e1291\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>. Additionally, a converted copy of SMN2 annotated as a duplication of SMN1 is enriched in African populations (F statistic\u2009=\u20090.43).<\/p>\n<p>We then measured the divergence of duplicated genes from their reference copies, indicating recent or ancient duplications and providing a measure of reference bias from missing paralogs. We constructed multiple-sequence alignments (MSAs; <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>) for sequences of each matrix and measured all pairwise differences in nonrepetitive sequences. We determined the average paralog sequence divergence relative to the ortholog divergence (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), which we refer to as the relative paralog divergence (RPD). We also measured copy number diversity using mean absolute error (MAE), indicating the CNV level among populations (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4d<\/a>). Based on RPD, using density-based spatial clustering of applications with noise<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 43\" title=\"Ester, M., Kriegel, H., Sander, J. &amp; Xu, X. A density-based algorithm for discovering clusters in large spatial databases with noise. In KDD&#x2019;96: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (eds Simoudis, E. et al.) 226&#x2013;231 (AAAI, 1996).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR43\" id=\"ref-link-section-d26018869e1317\" rel=\"nofollow noopener\" target=\"_blank\">43<\/a>, we identified two peaks at 0.71 and 3.2, with MAE centers at 0.18 and 0.93, corresponding to genes with rare and recent CNVs and more divergent and common CNVs, respectively. The latter reflect CNVs on different structural haplotypes that cannot be analyzed using a single reference genome. For example, AMY1A has a high RPD at 3.10 because of truncated duplications. These results are consistent with ancient bursts of duplications in human evolution<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Dennis, M. Y., et al. The evolution and population diversity of human-specific segmental duplications. Nat. Ecol. Evol. 1, 69 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR44\" id=\"ref-link-section-d26018869e1325\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a>.<\/p>\n<p>We next used ctyper genotypes to investigate recombination at different CNV loci. We determined multiallelic LD<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 45\" title=\"Okada, Y. eLD: entropy-based linkage disequilibrium index between multiallelic sites. Hum. Genome Var. 5, 29 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR45\" id=\"ref-link-section-d26018869e1332\" rel=\"nofollow noopener\" target=\"_blank\">45<\/a> (mLD; <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>) between PAs using the unrelated 1kGP samples for 989 subgroups that were adjacent less than 100\u2009kb apart in GRCh38 and reported the average mLD within each matrix (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4e<\/a>). There was a stronger negative rank correlation between MAE of copy number and mLD (\u03c1\u2009=\u2009\u22120.24, P value\u2009=\u20093.4\u2009\u00d7\u200910\u221215, Spearman\u2019s rank) than the rank correlation between mLDs and locus length (\u03c1\u2009=\u2009\u22120.21, P value\u2009=\u20091.5\u2009\u00d7\u200910\u221211, Spearman\u2019s rank), suggesting a reduced haplotype linkage in genes with frequent CNVs. The lowest mLD (0.013) was found in FAM90, a gene with frequent duplications and rearrangements<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 46\" title=\"Bosch, N. et al. Characterization and evolution of the novel gene family FAM90A in primates originated by multiple duplication and rearrangement events. Hum. Mol. Genet. 16, 2572&#x2013;2582 (2007).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR46\" id=\"ref-link-section-d26018869e1362\" rel=\"nofollow noopener\" target=\"_blank\">46<\/a>. The 29 loci with highest mLD (mLD\u2009&gt;\u20090.7) were enriched in the sex chromosomes (n\u2009=\u200919). Furthermore, HLA-B and HLA-DRB had mLD\u2009&gt;\u20090.7 and only deletion CNV (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Notes<\/a>).<\/p>\n<p>eQTL analysis<\/p>\n<p>To investigate the impact of paCNs on expression, we performed eQTL analysis using the Genetic European Variation in Disease<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Lappalainen, T. et al. Transcriptome and genome sequencing uncovers functional variation in humans. Nature 501, 506&#x2013;511 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR47\" id=\"ref-link-section-d26018869e1384\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a> (GEUVADIS) and GTEx<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 48\" title=\"GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318&#x2013;1330 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR48\" id=\"ref-link-section-d26018869e1388\" rel=\"nofollow noopener\" target=\"_blank\">48<\/a> cohorts. There were 4,512 genes that could be uniquely mapped in RNA-seq alignments. An additional 44 genes, such as SMN1, SMN2, AMY1A, AMY1B and AMY1C, have indistinguishable transcription products and were analyzed by pooling among all copies. We assigned PAs to these transcripts based on exonic sequences and performed association analysis with paCNs (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>).<\/p>\n<p>After merging paCNs to aggreCNs, 5.5% (178 of 3,224) of transcripts showed significance (corrected P\u2009=\u20091.6\u2009\u00d7\u200910\u22125, Pearson correlation) as previously observed<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Handsaker, R. E. et al. Large multiallelic copy number variations in humans. Nat. Genet. 47, 296&#x2013;303 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR31\" id=\"ref-link-section-d26018869e1422\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>. By contrast, when updating aggeCNs by individual paCNs and performing multivariable linear regression on expression (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>), there were significant improvements in fit for 27.6% (890 of 3,224) of transcripts (corrected P\u2009=\u20091.6\u2009\u00d7\u200910\u22125, one-tailed F-test; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5a<\/a>). To test whether the fit was explained by the nonuniform expression of different alleles of the same reference gene, we used a linear mixed model (LMM; <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>) to regress total expression to individual subgroups and estimate allele-specific expression and then compared these values to other subgroups of the same matrix that were assigned to the same reference gene (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). For subgroups within solvable matrices and more than ten samples, we found that 7.94% (150 of 1,890) of paralogs and 3.28% (546 of 16,628) of orthologs had significantly different expression levels (corrected with sample size\u2009=\u2009number of paralogs\u2009+\u2009orthologs, corrected P\u2009=\u20092.7\u2009\u00d7\u200910\u22126, \u03c72 test; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5b<\/a>). Overall, paralogs were found to have reduced expression (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5c<\/a>), consistent with previous findings for duplicated genes<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 49\" title=\"Lan, X. &amp; Pritchard, J. K. Coregulation of tandem duplicate genes slows evolution of subfunctionalization in mammals. Science 352, 1009&#x2013;1013 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR49\" id=\"ref-link-section-d26018869e1463\" rel=\"nofollow noopener\" target=\"_blank\">49<\/a>.<\/p>\n<p>Fig. 5: The impact of allele-specific copy number variation on gene expression.<a class=\"c-article-section__figure-link\" data-test=\"img-link\" data-track=\"click\" data-track-label=\"image\" data-track-action=\"view figure\" href=\"https:\/\/www.nature.com\/articles\/s41588-025-02346-4\/figures\/5\" rel=\"nofollow noopener\" target=\"_blank\"><img decoding=\"async\" aria-describedby=\"Fig5\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2025\/10\/41588_2025_2346_Fig5_HTML.png\" alt=\"figure 5\" loading=\"lazy\" width=\"685\" height=\"861\"\/><\/a><\/p>\n<p>a, Q\u2013Q plot of associations of (blue) aggreCNs with gene expression in GEUVADIS samples (red) and the improvement of allele-specific copy number over aggreCN. b, Comparative gene expression Q\u2013Q plots of orthologs (blue) and paralogs (red). c, Fold change effect size of all significant alternative expressions in b. Fold changes as well as P values are shown. Top: density of fold change effect size of orthologs and paralogs. d, Preferential tissue expression of orthologs and paralogs. e, Top: comparison of different models for explained expression variance (R2). Bottom: quantification of variance explained by different representations at different levels of CNV frequencies: full paCN genotypes, aggreCN and known eQTL variants (var.). f, Case study on SMN genes showing decreased gene expression in SMN-converted. The average expression level in PEER-corrected GEUVADIS samples (n\u2009=\u2009386) is shown under different copy numbers of SMN1 (n\u2009=\u2009741), SMN2 (n\u2009=\u2009569) and SMN-converted (n\u2009=\u200989). Transcript levels are the total coverage of all isoforms, and the exon 7 splicing level is measured by counting isoforms with a valid exon 7 splicing junction. g, Case study on amylase genes showing increased gene expression of translocated AMY2B using PEER-corrected GTEx pancreas data (no duplications, n\u2009=\u2009209; ordinary duplications, n\u2009=\u20096; AMY2B to AMY1, n\u2009=\u200925; AMY2B to AMY2A, n\u2009=\u20094; RNA-seq samples, n\u2009=\u2009304; box shows median and interquartile range; whiskers extend to 1.5\u00d7 interquartile range, with outliers beyond).<\/p>\n<p>We compared expression in 57 tissues in the GTEx samples to test for preferential expression of paralogs (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>). There was alternative tissue specificity for 132 of 2,820 paralogs (4.68%) and 225 of 19,197 orthologs (1.17%) (corrected P\u2009=\u20096.4\u2009\u00d7\u200910\u22128, union of two \u03c72 tests; <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> and Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5d<\/a>).<\/p>\n<p>Additionally, we used analysis of variance (ANOVA) to estimate the proportion of expression variance (R2) explained by paCNs in GEUVADIS expression data and compared it to that in a model based on known SNPs, indels and eQTL structural variants (SVs)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Keys, K. L., et al. On the cross-population generalizability of gene expression prediction models. PLoS Genet. 16, e1008927 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR50\" id=\"ref-link-section-d26018869e1602\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a> (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Sec10\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>). As expected, the highly granular paCNs explained the most variance: on average, 10.3% (14.3% including baseline). By contrast, 58.0% of transcripts are genes with known eQTL variants that explained valid variance by 2.14% (1.60% considering experimental noise, in agreement with a previous estimate of 1.97%<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Gamazon, E. R. et al. A gene-based association method for mapping traits using reference transcriptome data. Nat. Genet. 47, 1091&#x2013;1098 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR51\" id=\"ref-link-section-d26018869e1609\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>). On average, 1.98% of the variance was explained by aggreCNs, and 8.58% was explained by subgroup information. When combining both paCNs and known eQTL sites, 10.4% (19.0% including baseline) of the valid variance was explained (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5e<\/a>).<\/p>\n<p>We examined the SMN and AMY2B genes as case studies due to their importance in disease and evolution<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Lorson, C. L., Hahnen, E., Androphy, E. J. &amp; Wirth, B. A single nucleotide in the SMN gene regulates splicing and is responsible for spinal muscular atrophy. Proc. Natl Acad. Sci. USA96, 6307&#x2013;6311 (1999).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR27\" id=\"ref-link-section-d26018869e1623\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 52\" title=\"Pajic, P. Independent amylase gene copy number bursts correlate with dietary preferences in mammals. eLife 8, e44628 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR52\" id=\"ref-link-section-d26018869e1626\" rel=\"nofollow noopener\" target=\"_blank\">52<\/a>. The SMN genes were classified as SMN1, SMN2 and SMN-converted. We found no significant difference between the expression of all transcripts of SMN1 and SMN2 (0.281\u2009\u00b1\u20090.008 versus 0.309\u2009\u00b1\u20090.009; P\u2009=\u20090.078, \u03c72 test). However, significant differences were found between SMN-converted, and SMN1 and SMN2 (0.226\u2009\u00b1\u20090.012 versus 0.294\u2009\u00b1\u20090.002; P\u2009=\u20091.75\u2009\u00d7\u200910\u22127, \u03c72 test), with a 23.0% reduction in SMN-converted expression. By contrast, despite having lower overall expression, SMN-converted had 5.93\u00d7 the expression of valid exon 7 splicing<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 53\" title=\"Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15&#x2013;21 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR53\" id=\"ref-link-section-d26018869e1666\" rel=\"nofollow noopener\" target=\"_blank\">53<\/a> of SMN2 (P\u2009=\u20092.2\u2009\u00d7\u200910\u221216, \u03c72 test), indicating that SMN-converted has full functional splicing<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 54\" title=\"Lorson, C. L., Rindt, H. &amp; Shababi, M. Spinal muscular atrophy: mechanisms and therapeutic strategies. Hum. Mol. Genet. 19, R111&#x2013;R118 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#ref-CR54\" id=\"ref-link-section-d26018869e1683\" rel=\"nofollow noopener\" target=\"_blank\">54<\/a> but lower overall expression (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5f<\/a>). We studied the expression of AMY2B duplications, including alleles translocated proximally to other AMY genes, such as the PAs containing AMY1 and AMY2B in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a>. Using probabilistic estimation of expression residuals (PEER)-corrected GTEx pancreas data, we found that translocated AMY2B genes had significantly higher expression than other duplications (1.384\u2009\u00b1\u20090.233 versus \u22120.275\u2009\u00b1\u20090.183, P\u2009=\u20097.87\u2009\u00d7\u200910\u22129, \u03c72 test) (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-025-02346-4#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5g<\/a>).<\/p>\n","protected":false},"excerpt":{"rendered":"Overview of the genotyping method We represent variation as haplotype segments that are short enough to minimize disruption&hellip;\n","protected":false},"author":2,"featured_media":225520,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25],"tags":[3878,3880,3875,49,48,3877,7821,3879,3673,316,8785,3876,66,107921,2940],"class_list":["post-225519","post","type-post","status-publish","format-standard","has-post-thumbnail","category-genetics","tag-agriculture","tag-animal-genetics-and-genomics","tag-biomedicine","tag-ca","tag-canada","tag-cancer-research","tag-gene-expression","tag-gene-function","tag-general","tag-genetics","tag-genomics","tag-human-genetics","tag-science","tag-sequence-annotation","tag-software"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/225519","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=225519"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/225519\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/225520"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=225519"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=225519"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=225519"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}