{"id":681524,"date":"2026-07-09T02:40:21","date_gmt":"2026-07-09T02:40:21","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/681524\/"},"modified":"2026-07-09T02:40:21","modified_gmt":"2026-07-09T02:40:21","slug":"a-blended-genome-and-exome-sequencing-method-captures-genetic-variation-in-an-unbiased-and-cost-effective-manner","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/681524\/","title":{"rendered":"A blended genome and exome sequencing method captures genetic variation in an unbiased and cost-effective manner"},"content":{"rendered":"<p>Ethics approval and consent to participate<\/p>\n<p>All cohorts involving human participants that were included in this research have Institutional Review Board (IRB) approval following a full review of IRB applications and study consent forms by appropriate ethics committees. Details on each of the participating cohorts and IRB approvals are listed below:<\/p>\n<p>Hispanic ASD samples, IRB title: The Study of Novel Autism Genes, protocol no. 2012p001018, IRB of record: Mass General Brigham (previously known as PARTNERS), principal investigator on IRB: M. Daly.<\/p>\n<p>Chinese samples, IRB title: Neuropsychiatric genetics of a Chinese population from Shanghai Province, protocol no. IRB17-1379, IRB of record: Harvard T.H. Chan School of Public Health (HSPH), principal investigator on IRB: B. Neale.<\/p>\n<p>NeuroGAP-Psychosis:Ethiopia and South African samples, IRB title: Neuropsychiatric Genetics of African Populations \u2013 Psychosis (NeuroGAP-Psychosis), protocol no. IRB17- 0822, IRB of record: Harvard T. H. Chan School of Public Health (HSPH), principal investigator on IRB: K. Koenen.<\/p>\n<p>Paisa and Genomic Psychiatry Cohort, IRB title: Molecular Profiling of Psychiatric Disease, protocol no. 2014p001342, IRB of record: Mass General Brigham (previously known as Partners Healthcare), principal investigator on IRB: B. Neale.<\/p>\n<p>Inflammatory Bowel Disease samples, IRB title: The Broad Institute Study of Inflammatory Bowel Disease Genetics, protocol no. 2013P002634, IRB of record: Mass General Brigham (previously known as Partners Healthcare), principal investigator on IRB: R. Xavier.<\/p>\n<p>All participants who joined the studies listed under the above IRBs consented to the de-identified use of their data in disease research studies. During the consenting process, study staff explain the research goals and the basics of genetic processing activities. Participants are also given the option to opt out of studies if they choose. Samples used in this research and method development are all de-identified and have all been fully consented for inclusion in the study and any resulting publications.<\/p>\n<p>Wet lab protocolBiological samples<\/p>\n<p>BGE DNA samples have been extracted from saliva and whole blood specimens outside our laboratory. Samples derived from whole blood are generally preferred. Lower alignment rates of sequencing reads have been observed in saliva samples due to the presence of bacterial species. Furthermore, saliva samples can be more difficult to accurately quantify due to their consistency and heterogeneity. Following DNA extraction, we follow the BGE protocol described previously and shown in Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>.<\/p>\n<p>PCR-free library preparation<\/p>\n<p>After an initial quantification, DNA is normalized to 50\u2009ng\u2009\u03bcl\u22121 and transferred into a 384-well plate. Normalized DNA is purified using an automated 2.75\u00d7 solid-phase reversible immobilization (SPRI) clean-up with Ampure XP Agencourt beads (Beckman Coulter). Cleaned DNA is then quantified by spectrophotometry (Lunatic, Unchained Labs) and normalized to 25\u2009ng\u2009\u03bcl\u22121. DNA (target of 134\u2009ng input) undergoes a reduced and customized fragmentation\/end-repair\/A-tailing reaction for Illumina-compatible PCR-free library construction with custom NEBNext Ultra II FS DNA Library Preparation Kits (New England Biolabs) using the following conditions: 37\u2009\u00b0C for 42.57\u2009min, 65\u2009\u00b0C for 30\u2009min. Unique, dual-indexed adaptors (NEBNext Unique Dual Index UMI Adaptors, New England Biolabs) are ligated to fragments (20\u2009\u00b0C for 20\u2009min), and libraries undergo two consecutive SPRI size selections (0.5\u00d7 and 0.55\u00d7). Cleaned and size-selected PCR-free libraries are quantified by qPCR (Kapa Library Quantification Kit, Roche), then normalized (to ~0.3\u20132\u2009nM, depending on concentrations) before pooling into a single tube and concentrated (Unagi, Unchained Labs).<\/p>\n<p>The oligonucleotides used are commercially available. From New England Biolabs, we use four plates containing 384 unique adapters (cat. no. E7874L, E7876L, E7878L and E7395L), which also include the forward and reverse primer mix for exome capture.<\/p>\n<p>Exome-captured library preparation<\/p>\n<p>An aliquot from the prenormalized and prepooled PCR free libraries is used as input for PCR amplification (cycles are 98\u2009\u00b0C for 30\u2009s; 12 cycles of 98\u2009\u00b0C for 10\u2009sec and 65\u2009\u00b0C for 75\u2009s; 65\u2009\u00b0C for 5\u2009min) using the NEBNext Ultra II FS Library Preparation Kit and primers from the indexed adaptor kit (New England Biolabs). PCR-amplified libraries are quantified by spectrophotometer (Lunatic), purified with a 1\u00d7 SPRI-cleanup (Ampure) and normalized to 70\u2009ng\u2009\u03bcl\u22121. Samples are then pooled and undergo exome capture (Twist Alliance Clinical Research Exome probes from Twist Biosciences) using the recommended hybridization-capture protocol for xGen Hybridization Capture Core Reagents (Integrated DNA Technologies).<\/p>\n<p>Blending of PCR-free genome- and exome-captured libraries<\/p>\n<p>PCR-free and exome-captured pools are both qPCR-quantified on the same qPCR run. Nanomolar concentrations are taken into account to calculate the appropriate volumes to blend 33% WES with 67% WGS. BGE samples are again qPCR-quantified for sequencer loading calculations.<\/p>\n<p>Sequencing<\/p>\n<p>BGE blended pools with 384 samples containing unique barcodes are sequenced across six lanes of NovaSeqS4 (Illumina) with 2\u00d7 150-bp runs.<\/p>\n<p>Datasets analyzed to evaluate BGE quality<\/p>\n<p>BGE data used in these analyses were generated at the Broad Clinical Lab. Sample cohorts included in the dataset were recruited and submitted from the collaborating institutions for the PUMAS project, which includes cohorts from the GPC, Paisa population and NeuroGAP-Psychosis, as described below, in Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"table anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#Tab2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> and in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>.<\/p>\n<p>The GPC is a multi-institutional collaboration led by Rutgers University. The GPC resource includes a National Institute of Mental Health (NIMH)-managed repository of genomic samples at SAMPLED, genotypic and sequence data, and detailed clinical and demographic data for investigations of schizophrenia, bipolar disorder, obsessive\u2013compulsive disorder and coronavirus disease from a variety of ancestries collected in the USA. Within this analysis, we included 4,553 samples with 3,926 passing QC filters.<\/p>\n<p>The Paisa population is a genetic isolate from Colombia that has expanded rapidly following a series of migration-related bottlenecks. They have been the focus of genetics studies in neuropsychiatric disorders in the last decade. Paisa BGE data generated here emerged from a long-standing collaboration between teams at the Universidad de Antioquia, Medellin, Colombia, and University of California, Los Angeles, USA, to recruit samples from this population. Within the analysis, we included 9,007 samples, with 8,200 passing QC filters.<\/p>\n<p>The Stanley Center at the Broad Institute of MIT and Harvard initiated the NeuroGAP-Psychosis project in 2015<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 27\" title=\"Stevenson, A. et al. Neuropsychiatric Genetics of African Populations-Psychosis (NeuroGAP-Psychosis): a case&#x2013;control study protocol and GWAS in Ethiopia, Kenya, South Africa and Uganda. BMJ Open 9, e025469 (2019).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR27\" id=\"ref-link-section-d66665898e4342\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>, a collaboration with colleagues at Addis Ababa University in Ethiopia, KEMRI-Wellcome Trust in Kenya, Makerere University in Uganda, Moi University\/Moi Teaching and Referral Hospital in Kenya, the University of Cape Town in South Africa, and the Harvard T. H. Chan School of Public Health in the USA. With initial plans to recruit 35,000 participants (half cases with a diagnosis of schizophrenia or bipolar disorder and half controls), the target was expanded to 39,000 participants via the PUMAS Project awarded by NIMH. Sample recruitment ultimately exceeded the revised target, with &gt;42,000 samples collected across the five NeuroGAP-Psychosis collection sites. Within the analysis, we included 39,886 samples, with 35,138 passing QC filters.<\/p>\n<p>The Australian Genetics of Bipolar Study<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 19\" title=\"Lind, P. A. et al. Preliminary results from the Australian Genetics of Bipolar Disorder Study: a nation-wide cohort. Aust. N. Z. J. Psychiatry 57, 1428 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR19\" id=\"ref-link-section-d66665898e4350\" rel=\"nofollow noopener\" target=\"_blank\">19<\/a> led by the QIMR team is a national cohort of adults diagnosed with either type I or type II bipolar disorder. Genotypes were obtained using the Isohelix GeneFix GFX-02 2-ml saliva collection devices. The QSkin Sun and Health Study<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Olsen, C. M. et al. Cohort profile: the QSkin Sun and Health Study. Int. J. Epidemiol. 41, 929&#x2013;929i (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR20\" id=\"ref-link-section-d66665898e4354\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a> includes individuals from Queensland randomly sampled from the Australian Electoral Roll, with saliva samples for genotyping obtained using the Oragene DNA self-collection kit. Given the differing kits and enrollment strategies used in sample collection, we analyzed these cohorts separately for the imputation concordance analysis. Within the analysis and totaling from both QIMR cohorts, we included 7,499 samples, with 7,209 passing QC filters.<\/p>\n<p>Data from the PUMAS project are being deposited at the NIMH Data Archive (NDA). At the time of manuscript submission, genomic data from the first 10,000 samples have been submitted. All remaining samples from the PUMAS grant will be submitted at the end of the grant period. Data that are designated with NDA-GRU data use will be deposited into DNA collection #3805. Data designated with Disease-Specific (Mental Health) \u2013 DS (Mental Health) will be deposited into DNA collection #4538. Finally, data designated Health\/Medical\/Biomedical, NDA-HMB-MDS will be deposited into DNA collection #4539.<\/p>\n<p>Exome quality filtering<\/p>\n<p>We conducted QC filtering using Hail v0.28.128 to restrict to sites and variants with high confidence in exome data. We filtered out sites with more than six alleles that failed VQSR, were located in low-complexity regions, or fell outside the Twist target capture regions. Within an individual, genotype calls were filtered if read depth was &lt;10\u00d7, genotype quality was &lt;20, or allele balance was &lt;0.2 or &gt;0.8 in heterozygous calls, or &lt;0.8 in homozygous alternate calls.<\/p>\n<p>We filtered samples based on WES and WGS coverage, ancestry and WES sample quality metrics to restrict to high-quality samples for subsequent analysis. First, we removed samples with low WES or WGS coverage. Exome coverage per sample was determined by calculating the fraction of the exome target covered with at least 10\u00d7 read depth. Samples with exome fractions less than 90% or estimated WGS coverage of less than 1\u00d7 were removed. In addition, samples with chimeric or contamination read rates greater than 5% were removed, resulting in 853 removed samples.<\/p>\n<p>Next, we determined genetic ancestry of the PUMAS samples using the quality-filtered WES portion of the BGE. We combined the PUMAS data with four reference panels with diverse ancestries: HGDP, 1kGP, AWI-GEN and the African Genome Variation Project (AGVP)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Koenig, Z. et al. A harmonized public resource of deeply sequenced diverse human genomes. Genome Res. 34, 796&#x2013;809 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR18\" id=\"ref-link-section-d66665898e4375\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 28\" title=\"Gurdasani, D. et al. The African Genome Variation Project shapes medical genetics in Africa. Nature 517, 327&#x2013;332 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR28\" id=\"ref-link-section-d66665898e4378\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 29\" title=\"Ramsay, M. et al. H3Africa AWI-Gen Collaborative Centre: a resource to study the interplay between genomic and environmental risk factors for cardiometabolic diseases in four sub-Saharan African countries. Glob. Health Epidemiol. Genom. 1, e20 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR29\" id=\"ref-link-section-d66665898e4381\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>. Throughout this Article, we use ancestry labels assigned by existing genomic reference panels, including EUR (European), AFR (African) and AMR (Admixed American\u2014an imprecise label introduced by the 1000 Genomes Project to describe individuals with recent admixture from multiple continents including Indigenous American ancestry). Sites overlapping between the PUMAS data and all reference datasets were filtered to biallelic variants with call rates &gt;0.98 and MAFs &gt;0.1%. We conducted linkage disequilibrium (LD) pruning to extract independent markers and calculated the top ten principal components across all samples using Hail\u2019s hwe_normalized_pca function. We fit a random forest algorithm using the top ten principal components to the reference datasets of known ancestry and then applied the algorithm to the PUMAS data. We required PUMAS samples to have a probability of at least 0.7 to assign ancestry. Samples with probabilities less than 0.7 were removed (N removed = 1,799).<\/p>\n<p>We calculated sample quality metrics using Hail\u2019s sample_qc function. Within each ancestry and cohort combination, we removed samples with outlier values defined as values more than four median absolute deviations from the mean for the following metrics: N singletons, N insertions, N deletions, N transitions, N transversions, heterozygosity ratio, transition-to-transversion ratio and insertion-to-deletion ratio. Sample quality metrics filtering removed 2,317 samples.<\/p>\n<p>Finally, we removed samples with discrepancies between imputed genetic sex and reported gender. Genetic sex was imputed separately for each genetic ancestry group by calculating the inbreeding coefficient on the X chromosome using common, independent markers after removing markers within the pseudoautosomal region. Samples with an inbreeding coefficient &lt;0.6 were labeled \u2018female\u2019, and samples with an inbreeding coefficient &gt;0.6 were called \u2018male\u2019. We removed samples whose imputed sex and reported gender did not match, resulting in 316 removed samples.<\/p>\n<p>CNV calling and QCCNV calling<\/p>\n<p>To call CNVs from the deep-coverage exome data, GATK-gCNV 4.1.0.0 was used on the exome intervals following a previously described pipeline<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Babadi, M. et al. GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data. Nat. Genet. 55, 1589&#x2013;1597 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR10\" id=\"ref-link-section-d66665898e4423\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>. GATK-gCNV adjusts for known WES read-depth confounders such as GC content and mappability, while simultaneously adjusting for unspecified technical batch confounders such as sample extraction, sequencing, library preparation and mapping quality. We took the read-depth data as input from the BGE samples over a set of canonical transcript genomic intervals. In brief, the raw sequencing files were compressed into counts of the reads over the set of annotated exons and used as input into GATK-gCNV. A principal component analysis-based approach was used on the compressed observed counts to identify differences in the capture kit. Within each cohort, principal component analysis was used to define ancestrally similar batches of 1,000 samples; from each, a random subset of 200 samples were used to train a CNV-discovery model tailored to each batch. Subsequently, a distance-based and hybrid-density-based clustering approach was used to curate samples into batches to process in parallel. Once batching determination was completed, GATK-GCNV was used on each batch and metrics for filtering were produced by the Bayesian model underlying the data to balance recall and PPV.<\/p>\n<p>CNV benchmarking<\/p>\n<p>To benchmark BGE CNV exome data, we used data from 100 quads from the SSC<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Fu, J. M. et al. Rare coding variation provides insight into the genetic architecture and phenotypic context of autism. Nat. Genet. 54, 1320&#x2013;1331 (2022).\" href=\"#ref-CR30\" id=\"ref-link-section-d66665898e4435\">30<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Buxbaum, J. D. et al. The Autism Sequencing Consortium: large-scale, high-throughput sequencing in autism spectrum disorders. Neuron 76, 1052&#x2013;1056 (2012).\" href=\"#ref-CR31\" id=\"ref-link-section-d66665898e4435_1\">31<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"De Rubeis, S. et al. Synaptic, transcriptional and chromatin genes disrupted in autism. Nature 515, 209&#x2013;215 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR32\" id=\"ref-link-section-d66665898e4438\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>, sequenced previously. Benchmarking was carried out on all rare CNVs (frequency &lt;1%). The 100 quads have existing CNV calls<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Buxbaum, J. D. et al. The Autism Sequencing Consortium: large-scale, high-throughput sequencing in autism spectrum disorders. Neuron 76, 1052&#x2013;1056 (2012).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR31\" id=\"ref-link-section-d66665898e4442\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"De Rubeis, S. et al. Synaptic, transcriptional and chromatin genes disrupted in autism. Nature 515, 209&#x2013;215 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR32\" id=\"ref-link-section-d66665898e4445\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a> from deep WGS (~30\u00d7), and the CNV calls from WGS were used as a gold standard to benchmark the BGE CNVs. Recall was calculated by the proportion of WGS data sites that had a match in the BGE CNV callset. Moreover, for any given site, if at least 50% of samples that had that variant in the WGS data also had a GATK-gCNV call with consistent directionality (duplication or deletion) that overlapped at least 50% of captured intervals, this was considered a correct call. The optimal recall and PPV were observed for CNVs spanning more than four exons in the 100 quads. Samples were removed if there were more than ten rare CNVs, defined as &lt;1% frequency across the cohort. Any samples with &gt;more than ten high-quality CNVs were removed. We additionally removed samples that failed SNV QC. Subsequently, CNVs were filtered to have a quality score &gt;200 and restricted to &lt;1% frequency for each ancestry. We further benchmarked the performance of de novo CNVs derived from BGE data using GATK-gCNV as a function of the number of captured exons of canonical transcripts compared with validated WGS de novo CNVs from the quads. We additionally benchmarked the CNVs in saliva and blood against two independent constrained gene sets (top 1,000) based on LOEUF<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 12\" title=\"Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434&#x2013;443 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR12\" id=\"ref-link-section-d66665898e4449\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a> from gnomAD v2 and GISMO-mis<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 13\" title=\"Liao, C. et al. The landscape of gene loss and missense variation across the mammalian tree informs on gene essentiality. Preprint at bioRxiv &#010;                https:\/\/doi.org\/10.1101\/2024.05.16.594531&#010;                &#010;               (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR13\" id=\"ref-link-section-d66665898e4453\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>.<\/p>\n<p>Structural variant calling<\/p>\n<p>The structural variants were identified using VISTA v1<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Sarwal, V. et al. VISTA: an integrated framework for structural variant discovery. Brief. Bioinform. 25, bbae462 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR14\" id=\"ref-link-section-d66665898e4465\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a> for deletion and INSurVeyor v1.1.3<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Rajaby, R. et al. INSurVeyor: improving insertion calling from short read sequencing data. Nat. Commun. 14, 3243 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR15\" id=\"ref-link-section-d66665898e4469\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a> for insertions. VISTA is an ensemble tool that combines the best predictions from multiple individual callers, including Manta, Delly and others. We ran the individual callers and combined the results while applying the default parameters and quality filters. For InSurVeyor, we also applied the default parameters, which produced a high-quality callset that was used for analysis.<\/p>\n<p>Imputation<\/p>\n<p>To impute the low-coverage WGS data, we used the Genotype Likelihoods IMputation and Phasing (GLIMPSE2 v2.0.0) method<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 16\" title=\"Rubinacci, S., Hofmeister, R. J., Sousa da Mota, B. &amp; Delaneau, O. Imputation of low-coverage sequencing data from 150,119 UK Biobank genomes. Nat. Genet. 55, 1088&#x2013;1090 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR16\" id=\"ref-link-section-d66665898e4482\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>. We used the HGDP+1kGP reference panel<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Koenig, Z. et al. A harmonized public resource of deeply sequenced diverse human genomes. Genome Res. 34, 796&#x2013;809 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR18\" id=\"ref-link-section-d66665898e4486\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a> after filtering out singleton variants and indels, resulting in roughly about over 67 million variants available to be imputed.<\/p>\n<p>To phase and impute this large-scale data in a cost-effective manner, we used Broad\u2019s Hail Batch Service to submit randomized batches of 200 individuals. Then, we merged imputed results from the batches using bcftools. In total, we imputed over 55,000 samples for just under US$20,000, averaging about US$0.36 per sample (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>).<\/p>\n<p>AFs were corrected after merging by calculating a weighted mean of the batch allele frequencies (AFs) based on individual batch values (for i batches and N samples per batch):<\/p>\n<p>$${\\mathrm{AF}}_{\\mathrm{cohort}}=\\frac{\\sum {\\mathrm{AF}}_{i}{N}_{i}}{{\\sum N}_{i}\\,}.$$<\/p>\n<p>Similarly, INFO scores (I) were corrected post-merge via<\/p>\n<p>$${I}_{\\mathrm{cohort}}=1\\,-\\,\\frac{\\sum (1-{I}_{i})\\times 2{N}_{i}\\,\\times \\,{\\mathrm{AF}}_{i}(1-{\\mathrm{AF}}_{i})}{2\\times \\sum {N}_{i}\\times \\frac{\\sum {\\mathrm{AF}}_{i}{N}_{i}}{\\sum {N}_{i}}\\left(1-\\frac{\\sum {\\mathrm{AF}}_{i}{N}_{i}}{\\sum {N}_{i}}\\right)}.$$<\/p>\n<p>This process resulted in a final set of chromosome-specific BCF files, with all individuals of a cohort included and allele frequencies and INFO scores updated to reflect the full sample size.<\/p>\n<p>QC of GSA data<\/p>\n<p>Variants genotyped via the Illumina GSA platform were first filtered for a SNP call rate of at least 95%; then samples were filtered to have a call rate of at least 98%, heterogeneity F statistic or inbreeding coefficient |FHET| &lt;0.20, and concordance between genetically inferred and reported sex. A second SNP filtering step with a call rate of at least 98% was performed after the sample QC, followed by filtering steps to remove SNPs with missingness differences &gt;2% between cases and controls, and Hardy\u2013Weinberg equilibrium P value of 1\u2009\u00d7\u200910\u22126 for controls and 1\u2009\u00d7\u200910\u221210 for cases. We used King<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Manichaikul, A. et al. Robust relationship inference in genome-wide association studies. Bioinformatics 26, 2867&#x2013;2873 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR33\" id=\"ref-link-section-d66665898e4902\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a> to calculate the genetic relationship matrix and removed one sample in any pairs of individuals with second-degree or closer relatives. Finally, we used the conform-gt tool provided by the Beagle V5.4 imputation software<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Browning, B. L., Zhou, Y. &amp; Browning, S. R. A One-Penny imputed genome from next-generation reference panels. Am. J. Hum. Genet. 103, 338&#x2013;348 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR34\" id=\"ref-link-section-d66665898e4906\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a> to align the strand and allele order within the GSA VCFs to the HGDP+1kGP reference panel<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Koenig, Z. et al. A harmonized public resource of deeply sequenced diverse human genomes. Genome Res. 34, 796&#x2013;809 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR18\" id=\"ref-link-section-d66665898e4910\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>.<\/p>\n<p>Concordance between BGE and GSA GWAS array data<\/p>\n<p>To evaluate BGE data accuracy post-imputation, we compared BGE with previously generated Illumina GSA data as ground truth on the subset of individuals with both data types (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). Imputed variants were filtered to match those in the GSA. We computed nonreference concordance and aggregated R2 as a function of MAF for each cohort. To evaluate nonreference concordance<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Martin, A. R. et al. Low-coverage sequencing cost-effectively detects known and novel variation in underrepresented populations. Am. J. Hum. Genet. 108, 656&#x2013;668 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR3\" id=\"ref-link-section-d66665898e4929\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>, we define it as the number of true positives divided by the sum of the true positives, false negatives and false positives (excluding missing sites and SNPs with only reference alleles within the GSA subset), as shown in the formula below. We computed aggregated R2 per MAF bin by stacking all SNP dosages (from the imputed data) and genotype calls (from the GSA data) within a MAF bin, and calculating the squared Pearson\u2019s correlation coefficient. Note that, to avoid discrepancies in MAF for the NeuroGAP-Psychosis subsets (each with around 160 individuals), MAF values were based on the HGDP+1kGP AFR subset rather than in-sample allele frequencies. For the larger datasets, allele frequency values were taken from in-sample estimates.<\/p>\n<p>$$\\begin{array}{l}{\\text{non-ref-concordance}}({{\\text{SNP}}}_{j})\\\\=\\displaystyle\\frac{{\\rm{\\#}}\\,{\\rm{true\\; positives}}}{{\\rm{\\#}}\\,{\\rm{true\\; positives}}+{\\rm{\\#}}\\,{\\rm{false\\; positives}}+{\\rm{\\#}}\\,{\\rm{false\\; negatives}}}.\\end{array}$$<\/p>\n<p>Let XGSA be the matrix of genotypes for a given set of individuals and SNPs based on GSA array data, while XBGE, imputed is the matrix of imputed dosages for the same set of SNPs and individuals. We can calculate aggregated R2 per a given MAF bin for a SNP as follows<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 35\" title=\"The Genome of the Netherlands Consortium Whole-genome sequence variation, population structure and demographic history of the Dutch population. Nat. Genet. 46, 818&#x2013;825 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR35\" id=\"ref-link-section-d66665898e5123\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a>:<\/p>\n<p>$$\\begin{array}{l}\\mathrm{Aggregate}\\,{R}^{2}\\,\\mathrm{per}\\,\\mathrm{MAF}\\,\\mathrm{bin}\\\\={\\left({\\mathrm{Pearson}}^{^{\\prime} }{\\rm{s}}\\;\\mathrm{correlation}\\left(\\mathrm{vec}\\left({X}_{\\mathrm{GSA}}\\right),\\mathrm{vec}\\left({X}_{\\mathrm{BGE},\\mathrm{imputed}}\\right)\\right)\\right)}^{2}.\\end{array}$$<\/p>\n<p>Concordance between BGE and GSA by local ancestry inference<\/p>\n<p>We also evaluated post-imputation accuracy in BGE data stratified by each genotype\u2019s local ancestry background. We focus specifically on the Paisa and GPC populations, which have seen recent admixture between European, Native American Indigenous American and African continental ancestry. We first phased array genotypes via Beagle 5.4<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Browning, B. L., Tian, X., Zhou, Y. &amp; Browning, S. R. Fast two-stage phasing of large-scale sequence data. Am. J. Hum. Genet. 108, 1880&#x2013;1890 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR36\" id=\"ref-link-section-d66665898e5296\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a> using the full HGDP+1kGP reference panel (n\u2009=\u20094,099)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Koenig, Z. et al. A harmonized public resource of deeply sequenced diverse human genomes. Genome Res. 34, 796&#x2013;809 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR18\" id=\"ref-link-section-d66665898e5303\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>. Next, we followed the Tractor<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Atkinson, E. G. et al. Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nat. Genet. 53, 195&#x2013;204 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR37\" id=\"ref-link-section-d66665898e5307\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a> tutorial and computed local ancestry with RFMix2 software<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Maples, B. K., Gravel, S., Kenny, E. E. &amp; Bustamante, C. D. RFMix: a discriminative modeling approach for rapid and robust local-ancestry inference. Am. J. Hum. Genet. 93, 278&#x2013;288 (2013).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR22\" id=\"ref-link-section-d66665898e5311\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>. For the Paisa cohort, we modeled three-way local ancestry to capture admixture represented by EUR, AFR and AMR (as a proxy for Indigenous American ancestry) populations in the HGDP+1kGP reference panel. Specifically, we filtered the reference panel to only samples with AFR\/EUR\/AMR population labels. For local ancestry inference to be most useful, reference samples themselves should be representative of ancestral populations and non-admixed. Because AMR samples tend to themselves be recently admixed, we excluded 358\/549 reference AMR samples that have &lt;90% Indigenous American ancestry as measured by RFMix2. As a result, the total number of reference samples Ntotal\u2009=\u20091,935 are split among 191 AMR, 752 EUR and 992 AFR. For the GPC population, we inferred two-way local ancestry to capture European and African admixture. When running RFMix2, we used the optional EM flag set to one iteration to account for residual admixture within the subsetted reference panel. Finally, we also included the \u2018-n 5\u2019 flag to account for reference panel sample size imbalance.<\/p>\n<p>After local ancestry inference, two metrics were used to evaluate imputation accuracy for different ancestry backgrounds, including aggregate R2 and aggregate nonreference concordance, computed by stacking all SNPs within a MAF bin and evaluating a single nonreference concordance value. This is analogous to the aggregate R2 formula except that we compute concordance rather than squared correlation after stacking SNPs column-wise. Note that we favor the aggregate nonreference concordance metric as opposed to averaging SNP-by-SNP concordances due to the presence of rare SNPs. This is because a rare SNP may have only several copies of the minor allele, and by chance, they may all reside on one specific ancestry background. This leaves no copies of minor alleles for other ancestry backgrounds. Averaging SNP-by-SNP nonreference concordances will therefore include excess zeros for the remaining ancestries, artificially deflating their accuracy. Stacking rare SNPs rescues concordance for rare SNPs among all ancestry backgrounds by virtue of including more copies of the minor allele, ultimately producing more accurate and interpretable results.<\/p>\n<p>Number of variants per data generation type<\/p>\n<p>We evaluated a subset of samples from the NeuroGAP-Psychosis data (N\u2009=\u200979) that were assayed across three different data generation strategies, including BGE, Illumina GSA arrays and 30\u00d7 WGS. To count the number of variants per MAF category in the noncoding BGE data, we used the imputed data imputed with GLIMPSE2 after filtering for an INFO score of at least 0.8 and removing monomorphic variants. For the coding regions, we utilized the high-coverage exome data rather than the imputed data, filtering for VQSR PASS variants as well as excluding locus control regions and filtering for monomorphic variants. We required genotypes to have genotype quality &gt;50 and depth &gt;13, and normalized Phred-scaled likelihoods to be &gt;50. These same filters were applied for both coding and noncoding regions in the WGS dataset. For the Illumina GSA GWAS array data, we used the imputed data generated previously<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 3\" title=\"Martin, A. R. et al. Low-coverage sequencing cost-effectively detects known and novel variation in underrepresented populations. Am. J. Hum. Genet. 108, 656&#x2013;668 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#ref-CR3\" id=\"ref-link-section-d66665898e5342\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a> using the phase 3 1000 Genomes Project reference panel with BEAGLE v5.1 for all categories, filtering for dosage R2 of at least 0.8 and removing monomorphic variants. Coding regions were defined by Twist target capture region intervals.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02669-w#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Ethics approval and consent to participate All cohorts involving human participants that were included in this research have&hellip;\n","protected":false},"author":2,"featured_media":681525,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[5083,5085,3251,5082,8815,59,5084,3250,102,5081,13781,56,54,55],"class_list":["post-681524","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-agriculture","tag-animal-genetics-and-genomics","tag-biomedicine","tag-cancer-research","tag-dna-sequencing","tag-gb","tag-gene-function","tag-general","tag-health","tag-human-genetics","tag-population-genetics","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/681524","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=681524"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/681524\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/681525"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=681524"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=681524"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=681524"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}