{"id":793024,"date":"2026-10-05T21:10:14","date_gmt":"2026-10-05T21:10:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/793024\/"},"modified":"2026-10-05T21:10:14","modified_gmt":"2026-10-05T21:10:14","slug":"imputation-of-fluid-intelligence-scores-reduces-ascertainment-bias-and-increases-power-for-analyses-of-common-and-rare-variants","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/793024\/","title":{"rendered":"Imputation of fluid intelligence scores reduces ascertainment bias and increases power for analyses of common and rare variants"},"content":{"rendered":"<p>Ethics oversight<\/p>\n<p>UK Biobank obtained ethical approval from the NHS North West Centre for Research Ethics Committee (reference 11\/NW\/0382) and approved the use of data for this study.<\/p>\n<p>Sample selection<\/p>\n<p>We conducted all analyses only in individuals who genetically cluster with European ancestry (n\u2009=\u2009455,943). This was inferred by first projecting principal components (PCs) from 1000 Genomes onto the UKB participants and excluding participants who did not cluster with the European populations from the 1000 Genomes Project. After this, another principal component analysis was conducted to capture ancestry differences within the genetically more homogeneous individuals with European\/British ancestries<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Abdellaoui, A., Dolan, C. V., Verweij, K. J. H. &amp; Nivard, M. G. Gene-environment correlations across geographic regions affect genome-wide association studies. Nat. Genet. 54, 1345&#x2013;1354 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR72\" id=\"ref-link-section-d310663722e3156\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a>.<\/p>\n<p>FIS score transformations<\/p>\n<p>To correct for age decline in FI over time and differences in scales across FI tests, we transformed the FIS measures. For each FIS separately, we ran a regression model: FIS\u2009~\u2009age\u2009+\u2009age2, extracted the residuals and added the intercept to allow for differences in means between measures. The age used was the participant\u2019s age at the time of the respective measurement, which was either obtained from UKB variable \u2018age attended assessment center\u2019 (Data-Field 21003) or approximated from \u2018when FI test completed\u2019 (Data-Field 20135) using the participants\u2019 birth year and month.<\/p>\n<p>For the online measures (FIS4 and FIS5), we did an additional transformation, setting all scores of 14 to 13, before running the regression model. This broadly aligns the scales of the online measures (14 questions) with those of the in-person test (13 questions; Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1<\/a>).<\/p>\n<p>ImputationSoftImpute<\/p>\n<p>The R package SoftImpute<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 73\" title=\"Mazumder, R., Hastie, T. &amp; Tibshirani, R. Spectral regularization algorithms for learning large incomplete matrices. J. Mach. Learn. Res. 11, 2287&#x2013;2322 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR73\" id=\"ref-link-section-d310663722e3188\" rel=\"nofollow noopener\" target=\"_blank\">73<\/a> was used to impute FIS. SoftImpute is a matrix completion algorithm that approximates missing values by identifying and leveraging patterns in the available data. It does so by minimizing an objective function consisting of two terms\u2014(1) the distance between the observed entries in the observed and imputed data matrices according to the Frobenius norm, and (2) the product of a tuning parameter \u03bb and the sum of the singular values of the imputed data matrix. Intuitively, this means the algorithm fills in missing values in a way that minimizes the rank of the imputed data matrix, favoring low-dimensional structure over complexity. It starts with an initial guess for the missing values, then iteratively refines this guess by applying a soft-thresholded singular value decomposition on the complete matrix. The main parameters specified in SoftImpute are rank and \u03bb. Rank determines the maximum number of factors the method uses to represent the data, and should always be set to at most Nvariables \u2212 1. \u03bb controls the amount of smoothing applied during imputation. A high \u03bb aims at lower model complexity, making it less sensitive to noise and more prone to overlook nuances. Ideally, \u03bb is set to be slightly less than the rank set.<\/p>\n<p>Variable selection<\/p>\n<p>Our imputation strategy underwent several rounds of refinement (Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3<\/a>), one of which concerned the selection of imputation variables. Here we describe the criteria used for the initial selection and the steps taken to refine this set.<\/p>\n<p>                  Initial selection<\/p>\n<p>We first selected 152 UKB phenotypes among those collected at the initial assessment center visit that showed at least nominally significant correlation (P\u2009&lt;\u20090.05) with observed FI, focusing on FIS1 as it had the largest sample available at the time. For continuous phenotypes, we chose those with an absolute phenotype correlation (|rpheno|) with FIS1 &gt;0.05 at a significance threshold of P\u2009&lt;\u20090.05. Categorical phenotypes were first classified into ordinal and nominal types and then further evaluated. For ordinal phenotypes, we verified whether the values could be interpreted as a quantitative variable, reordered them if necessary and then retained those with |rpheno|\u2009&gt;\u20090.05. We converted nominal phenotypes with n categories into n\u2009\u2212\u20091 binary variables, after which we calculated the rpheno between each of those binary variables and FIS1. We only included variables with two or more categories having |rpheno|\u2009&gt;\u20090.1. We excluded phenotypes that are components of any FIS measure. In total, we selected 90 continuous phenotypes, 59 ordinal phenotypes and 3 nominal phenotypes (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). For the three nominal phenotypes, we selected the categories showing a strong correlation with FIS1 (|rpheno|\u2009&gt;\u20090.1), coded them as additional binary phenotypes and removed the original phenotypes, ultimately leading to a set of 154 variables used as our initial set. Finally, we coded all missing values (for example, \u2018preferred not to answer\u2019 or \u2018unknown\u2019) as NA. Based on the 154 resulting phenotypes, we ran imputation with SoftImpute parameters\u2014rank\u2009=\u2009150, \u03bb\u2009=\u2009120.<\/p>\n<p>                Final selection<\/p>\n<p>In our final selection, we narrowed the initial selection of variables down to phenotypes that correlate more specifically with cognitive signal. To examine which phenotypes fall within this criterion, we derived a GWAS of the \u2018noncognitive component of imputed FIS\u2019 (NonCog-iFIS) from our first FIS imputation. To do this, we applied a GenomicSEM model (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>) for GWAS-by-subtraction as applied in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Demange, P. A. et al. Investigating the genetic architecture of noncognitive skills using GWAS-by-subtraction. Nat. Genet. 53, 35&#x2013;44 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR30\" id=\"ref-link-section-d310663722e3282\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>. GenomicSEM is an R package that allows fitting structural equation models on summary statistics from GWAS. In GWAS-by-subtraction models, a model is fitted that includes two phenotypes (GWAS) and two latent factors. The first latent factor represents the commonalities between phenotypes by regressing both phenotypes on this latent factor. The second latent factor comprises the genetic variance unique to one of the phenotypes. This is achieved by regressing the remaining genetic variance for the phenotypes on the latent factor after regressing out the variance captured in the first latent factor. Subsequently, both latent factors can be regressed on individual SNPs to obtain GWAS. In our analysis, we ran a GWAS-by-subtraction model using intelligence as mentioned in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Savage, J. E. et al. Genome-wide association meta-analysis in 269,867 individuals identifies new genetic and functional links to intelligence. Nat. Genet. 50, 912&#x2013;919 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR14\" id=\"ref-link-section-d310663722e3286\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a>, and in our first iteration, imputed FIS values as phenotypes. The model allows us to capture the genetic variance unique to this imputed FIS in the \u2018NonCog\u2019 latent factor and subsequently regress individual SNPs on this factor to obtain the NonCog-iFIS GWAS.<\/p>\n<p>We then computed genetic correlations between each of the 154 imputation variables and intelligence in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 14\" title=\"Savage, J. E. et al. Genome-wide association meta-analysis in 269,867 individuals identifies new genetic and functional links to intelligence. Nat. Genet. 50, 912&#x2013;919 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR14\" id=\"ref-link-section-d310663722e3293\" rel=\"nofollow noopener\" target=\"_blank\">14<\/a> and the derived NonCog-iFIS GWAS (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>). We retained phenotypes in the final selection if the absolute genetic correlation was stronger with intelligence than with NonCog-iFIS (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>), resulting in 82 phenotypes being selected.<\/p>\n<p>For imputations using the final selection of variables, we adjusted our SoftImpute parameters (rank\u2009=\u200980, \u03bb\u2009=\u200970) due to the reduced number of selected phenotypes. We also explored the effects of different imputation parameters on imputation accuracy, but found that rank\u2009=\u200980 and \u03bb\u2009=\u200970 achieved the highest imputation accuracy (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>, left).<\/p>\n<p>Removing outliers<\/p>\n<p>After each imputation, we identified and removed outliers using the criterion:<\/p>\n<p>$${X_{i}}\\notin \\left(\\overline{X}-3\\sigma ,\\,\\overline{X}+3\\sigma \\right)$$<\/p>\n<p>where Xi represents an imputed score, \\(\\overline{X}\\) represents the mean of the measured score and \u03c3 represents the s.d. of the measured score.<\/p>\n<p>In other words, we removed individuals whose imputed score was more than 3 s.d. from the mean of the observed scores. This results in slightly differing sample sizes across the imputation approaches (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>).<\/p>\n<p>Evaluating accuracy<\/p>\n<p>To evaluate the imputation accuracy, we randomly selected 50,000 participants as the evaluation set. We introduced \u2018synthetic missingness\u2019 by setting FIS measures for these participants to be missing. After imputation, we examined the Pearson correlation between imputed FIS values and the original measures for our evaluation set and defined that correlation as the imputation accuracy. Only participants who have the targeted FI measure were included in calculation of accuracy for each FI measure (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a>).<\/p>\n<p>We also examined whether the imputation was more accurate in different parts of the phenotype distribution. We classified observed values of the evaluation set into low, medium and high terciles, then examined the imputation accuracy within each tercile using the same strategy as above. We found the imputation accuracy to be higher in the low and high terciles than in the medium tercile (Supplementary Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">9<\/a>, right).<\/p>\n<p>Combining measured and imputed FIS<\/p>\n<p>Combining imputed and measured FIS was done by mega-analysis in all approaches. Before mega-analysis, imputed and measured values were scaled separately to have a mean of 0 and standard deviation of 1. While this may in principle bias associations, the resulting per-SNP bias scales with allele frequency differences between measured and imputed individuals, which are likely negligible (see Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>\u2014Mega-analysis recovers population effects under calibrated imputation (Remark 10) for further discussion). We also evaluated combining imputed and measured values through meta-analysis (Supplementary Note <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>).<\/p>\n<p>GWAS<\/p>\n<p>For the GWASs, we used REGENIE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Mbatchou, J. et al. Computationally efficient whole-genome regression for quantitative and binary traits. Nat. Genet. 53, 1097&#x2013;1103 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR34\" id=\"ref-link-section-d310663722e3473\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a> (v4.1) step 1 and step 2 on the UKB research analysis platform. For step 1, we used genotyped SNPs from UKB Data-Field 22418 passing standard quality control (missingness\u2009\u2264\u20090.1, MAF\u2009\u2265\u20090.01, Hardy\u2013Weinberg equilibrium P\u2009\u2265\u20091\u2009\u00d7\u200910\u221215). Individuals with genotype missingness &gt; 0.1 were excluded.<\/p>\n<p>In step 2, we analyzed imputed SNPs from UKB Data-Field 22828 that are in the Haplotype Reference Consortium<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 74\" title=\"McCarthy, S. et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat. Genet. 48, 1279&#x2013;1283 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR74\" id=\"ref-link-section-d310663722e3485\" rel=\"nofollow noopener\" target=\"_blank\">74<\/a> and passing quality control in unrelated European individuals (missingness\u2009&lt;\u20090.05, MAF\u2009&gt;\u20090.001, Hardy\u2013Weinberg equilibrium P\u2009&gt;\u20091\u2009\u00d7\u200910\u221210). To obtain the heteroskedasticity-robust standard error estimator in REGENIE, we specified a dummy interaction covariate and set \u2018&#8211;rare-mac 0\u2019 to force the robust estimator to be used for all SNPs. The dummy interaction covariate was generated to be a random standard normal variable.<\/p>\n<p>The covariates (\u2018&#8211;covarFile\u2019) included in both steps were 25 genetic PCs that capture ancestry differences within European individuals<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Abdellaoui, A., Dolan, C. V., Verweij, K. J. H. &amp; Nivard, M. G. Gene-environment correlations across geographic regions affect genome-wide association studies. Nat. Genet. 54, 1345&#x2013;1354 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR72\" id=\"ref-link-section-d310663722e3500\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a>, age (at time of respective FIS measurement), age2, age\u2009\u00d7\u2009sex, age2\u2009\u00d7\u2009sex and the array used to measure each individual\u2019s genotype and sex as binary variables. For the average FIS approach, we computed the average age across measures and used that as the age covariate. For individuals with imputed average FIS, age was set to the UKB variable \u2018age initial assessment visit\u2019 (Data-Field 21003, Instance 0).<\/p>\n<p>Meta-analysis<\/p>\n<p>For meta-analyses, we used METAL<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 75\" title=\"Willer, C. J., Li, Y. &amp; Abecasis, G. R. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics 26, 2190&#x2013;2191 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR75\" id=\"ref-link-section-d310663722e3516\" rel=\"nofollow noopener\" target=\"_blank\">75<\/a> with the \u2018STDERR\u2019 approach. This weights effect size estimates by the inverse of corresponding standard errors. We only include SNPs with n\u2009&gt;\u200910,000.<\/p>\n<p>Genetic correlations and SNP h<br \/>\n                        2<\/p>\n<p>LDSC<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 71\" title=\"Bulik-Sullivan, B. et al. An atlas of genetic correlations across human diseases and traits. Nat. Genet. 47, 1236&#x2013;1241 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR71\" id=\"ref-link-section-d310663722e3537\" rel=\"nofollow noopener\" target=\"_blank\">71<\/a> was used to compute genetic correlations and SNP heritabilities. This tool requires munged (parsed) sumstats. In munging, we aligned SNPs with those in the HapMap 3 (ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 76\" title=\"International HapMap 3 Consortium. Integrating common and rare genetic variation in diverse human populations. Nature 467, 52&#x2013;58 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR76\" id=\"ref-link-section-d310663722e3541\" rel=\"nofollow noopener\" target=\"_blank\">76<\/a>) set using the \u2018&#8211;merge-alleles\u2019 flag. Heritabilities and genetic correlations were computed using the \u2018&#8211;h2\u2019 and \u2018&#8211;rg\u2019 flag, respectively, with default parameters. To estimate the genome-wide correlation between the direct effects and NTCs, we used the SNIPAR package correlate.py script<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Young, A. I. et al. Mendelian imputation of parental genotypes improves estimates of direct genetic effects. Nat. Genet. 54, 897&#x2013;905 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR33\" id=\"ref-link-section-d310663722e3545\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>. We applied a block-jackknife procedure to test whether estimates differed significantly across imputation approaches. To do this, we derived the intersection of SNPs included in each GWAS, ordered them by chromosome and then by position, and divided them into 200 blocks of ~5,077 SNPs.<\/p>\n<p>Identifying lead SNPs<\/p>\n<p>The online platform for FUMA<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Watanabe, K., Taskesen, E., van Bochoven, A. &amp; Posthuma, D. Functional mapping and annotation of genetic associations with FUMA. Nat. Commun. 8, 1826 (2017).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR32\" id=\"ref-link-section-d310663722e3557\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a> SNP2GENE was used to identify lead SNPs. First, independent significant SNPs are identified (P\u2009&lt;\u20095\u2009\u00d7\u200910\u22128, r2\u2009&lt;\u20090.6). Subsequently, identified significant SNPs are designated lead SNPs if they are independent from each other at a second threshold of r2\u2009&lt;\u20090.1.<\/p>\n<p>Gene prioritization and tissue expression<\/p>\n<p>MAGMA<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 77\" title=\"De Leeuw, C. A., Mooij, J. M., Heskes, T. &amp; Posthuma, D. MAGMA: generalized gene-set analysis of GWAS data. PLoS Comput. Biol. 11, e1004219 (2015).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR77\" id=\"ref-link-section-d310663722e3583\" rel=\"nofollow noopener\" target=\"_blank\">77<\/a> (as implemented in FUMA\u2019s SNP2GENE process) was used for gene prioritization and tissue expression analyses using the results from the population-based GWAS. MAGMA aggregates SNP-level data into gene-level data and performs a gene-based association test to identify significantly associated genes. The gene window was kept at 0\u2009kb, restricting analysis to SNPs located within a gene. Resulting gene-based P values were downloaded and FDR corrected using the Benjamini\u2013Hochberg procedure. To obtain our set of prioritized genes, we selected protein-coding genes with an FDR\u2009&lt;\u20091% that were also MANE select transcripts<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 78\" title=\"Morales, J. et al. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature 604, 310&#x2013;315 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR78\" id=\"ref-link-section-d310663722e3590\" rel=\"nofollow noopener\" target=\"_blank\">78<\/a>. We then applied the SNP2GENE process in FUMA to test whether the significant genes were enriched in particular tissues; this takes the gene-level P values from MAGMA as input. We used expression data from 54 tissues from GTEx (v8; ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 79\" title=\"GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 369, 1318&#x2013;1330 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR79\" id=\"ref-link-section-d310663722e3597\" rel=\"nofollow noopener\" target=\"_blank\">79<\/a>) as reference data.<\/p>\n<p>Within-family GWAS<\/p>\n<p>We conducted within-family GWAS in UKB using the SNIPAR package<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Young, A. I. et al. Mendelian imputation of parental genotypes improves estimates of direct genetic effects. Nat. Genet. 54, 897&#x2013;905 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR33\" id=\"ref-link-section-d310663722e3610\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a> in individuals with European ancestry. SNIPAR leverages the presence of genotyped first-degree relatives to impute missing parental genotypes, allowing for their downstream use to conduct within-family GWAS. We estimated pairwise kinship coefficients using KING<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 80\" title=\"Manichaikul, A. et al. Robust relationship inference in genome-wide association studies. Bioinformatics 26, 2867&#x2013;2873 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR80\" id=\"ref-link-section-d310663722e3614\" rel=\"nofollow noopener\" target=\"_blank\">80<\/a> and used default parameters to conduct the within-family GWAS using scripts provided in SNIPAR. We controlled for the same covariates as in the population GWAS. SNP heritability estimates were estimated using LDSC as described above. We filtered to SNPs with INFO score\u2009&gt;\u20090.98 for all analyses using SNIPAR-imputed individuals, including the GWAS and PGI analyses.<\/p>\n<p>Whole-exome sequencing data in UKB and quality control<\/p>\n<p>Whole-exome-sequencing data were generated at the Regeneron Genetics Center, and the sequencing procedure has been described in previous studies. We used custom applets to perform quality control for the whole-exome sequencing data of 469,836 participants within the UKB research analysis platform. First, we used BCFtools norm to split and left-align multi-allelic variants in the population-level Variant Call Format files into separate alleles. Next, we performed genotype-level filtering using BCFtools filter separately for single nucleotide variants (SNVs) and insertions\/deletions. Specifically, SNV genotypes with a depth below 7 and genotype quality below 20, or insertions\/deletion genotypes with a depth below 10 and genotype quality below 20, were set to missing. We also applied a binomial test to check for an expected alternate allele contribution of 50% for heterozygous SNVs, and SNV genotypes with a binomial test P\u2009\u2264\u20090.0001 were set to missing. Finally, we recalculated the proportion of missing genotypes for each variant and excluded all variants with more than 50% missingness.<\/p>\n<p>Next, we annotated the variants using the ENSEMBL Variant Effect Predictor (VEP; v104) with the &#8211;everything flag. For each variant, we prioritized a single ENSEMBL transcript based on whether the transcript was protein coding, MANE Select (v0.97) or the VEP canonical transcript. The variant consequence was prioritized based on severity as defined by VEP. After annotation, we grouped stop-gained, frameshift, splice acceptor and splice donor variants into a single PTV category. Missense and synonymous variant consequences were defined according to VEP criteria, and only autosomal variants within ENSEMBL protein-coding transcripts were retained for further analysis.<\/p>\n<p>We further filtered variants by MAF, retaining only those with MAF &lt;0.001%. These variants were annotated with LOFTEE, REVEL, AlphaMissense, and MPC for further filtering. For PTVs, only high-confidence PTVs defined by LOFTEE were retained. For missense variants, a damaging missense variant set was created by including variants with AlphaMissense scores &gt; 0.56, REVEL scores &gt; 0.5 and MPC scores &gt; 2.<\/p>\n<p>Exome-wide burden tests of rare coding variants<\/p>\n<p>To examine the association between FIS and the burden of rare coding variants, we counted the number of rare PTVs, missense variants, damaging missense variants (as described above) and synonymous variants both in exome-wide and in high loss-of-function-intolerant (pLI\u2009&gt;\u20090.9) genes. The variant burden was then used as the predictor variable in linear regression models, with FIS (observed, imputed and combined) as the response variables. The models were run in unrelated participants with European ancestry (n\u2009=\u2009328,795). We controlled for the top 25 PCs, as well as age, sex, age2 and the interactions between age and sex, and age2 and sex<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 72\" title=\"Abdellaoui, A., Dolan, C. V., Verweij, K. J. H. &amp; Nivard, M. G. Gene-environment correlations across geographic regions affect genome-wide association studies. Nat. Genet. 54, 1345&#x2013;1354 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR72\" id=\"ref-link-section-d310663722e3650\" rel=\"nofollow noopener\" target=\"_blank\">72<\/a>.<\/p>\n<p>Gene-based burden test of rare coding variants<\/p>\n<p>We performed gene-based burden tests on observed and combined FIS (standardized to mean 0 and variance 1) and EA (number of years of education) using a two-step regression analysis in REGENIE. As REGENIE accounts for relatedness and population structure, the gene-based tests were performed in all participants with European ancestry (n\u2009=\u2009438,285; please note that this is lower than the sample size used for common-variant analyses, as not all participants had exome data). In the first step, REGENIE fits a stacked block ridge regression to produce a leave-one-chromosome-out genetic prediction of the focal phenotype. The association test is then carried out in the second step by fitting regression models conditioned on the leave-one-chromosome-out predictions. For both steps, we adjusted for sex, age, age2, sex-by-age interaction, sex-by-age2 interaction, the top 25 PCs and recruitment centers (as categorical variables) to control for population structure and age. The FIS scores were rank based and inverse-normal transformed, as recommended by REGENIE. EA was coded in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 81\" title=\"Okbay, A. et al. Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nat. Genet. 54, 437&#x2013;449 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR81\" id=\"ref-link-section-d310663722e3669\" rel=\"nofollow noopener\" target=\"_blank\">81<\/a> by using individual reports of highest attained qualification, with \u2018college or university degree\u2019 counting as 20 years, \u2018other professional qualifications\u2019 as 15, \u2018A levels\/AS levels or equivalent\u2019 as 13, \u2018O levels\/GCSEs or equivalent\u2019, \u2018CSEs or equivalent\u2019 as 10 and \u2018none of the above\u2019 as 7. We ran both burden tests and SKAT-O tests on three consequence classes\u2014PTV, damaging missense and PTV\u2009+\u2009damaging missense. Thus, for FIS we conducted 12 tests per gene (including 2 phenotypes, 2 statistical methods, 3 consequence classes). To account for multiple testing, we calculated the FDR using the Benjamini\u2013Hochberg method across the vector of all P values and considered genes passing FDR\u2009&lt;\u20091% as \u2018significant\u2019. For EA, we conducted six tests per gene as we only considered a single phenotype, and similarly accounted for multiple testing by calculating the FDR across a vector of all P values and considering genes passing FDR\u2009&lt;\u20091% as \u2018significant\u2019. We used the DDG2P gene list downloaded on 5 February 2025 for annotating genes as \u2018well-established\u2019 developmental condition genes.<\/p>\n<p>Replication analyses in external cohorts<\/p>\n<p>Information about the replication cohorts (ALSPAC, MCS and INTERVAL) is given in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>. Quality control and imputation of genotype data in ALSPAC and MCS were conducted in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Malawsky, D. S. et al. Common and rare genetic variant associations with cognitive performance across development in British birth cohorts. Nat. Hum. Behav. 10, 1841&#x2013;1854 &#010;                https:\/\/doi.org\/10.1038\/s41562-026-02491-8&#010;                &#010;               (2026).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR82\" id=\"ref-link-section-d310663722e3691\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> and are summarized in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>, as is the preparation of the exome data, which was described in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 37\" title=\"Koko, M. et al. Exome sequencing of UK birth cohorts. Wellcome Open Res. 9, 390 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR37\" id=\"ref-link-section-d310663722e3698\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>. Quality control in INTERVAL was conducted in refs. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 83\" title=\"Sun, B. B. et al. Genomic atlas of the human plasma proteome. Nature 558, 73&#x2013;79 (2018).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR83\" id=\"ref-link-section-d310663722e3702\" rel=\"nofollow noopener\" target=\"_blank\">83<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 84\" title=\"Astle, W. J. et al. The allelic landscape of human blood cell trait variation and links to common complex disease. Cell 167, 1415&#x2013;1429 (2016).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR84\" id=\"ref-link-section-d310663722e3705\" rel=\"nofollow noopener\" target=\"_blank\">84<\/a> and is summarized in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>.<\/p>\n<p>Cognitive performance measures<\/p>\n<p>In ALSPAC, we considered IQ measured at age 8 using the Wechsler Intelligence Scale for Children test<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 85\" title=\"Wechsler, D. &amp; Kodama, H. in Wechsler Intelligence Scale for Children Vol. 1 (Psychological Corp, 1949).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR85\" id=\"ref-link-section-d310663722e3720\" rel=\"nofollow noopener\" target=\"_blank\">85<\/a>. We included 5,283 unrelated children with genetically inferred European ancestry and at least one genotyped parent in our analyses. In MCS, we derived a cognitive performance measure as previously described in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Malawsky, D. S. et al. Common and rare genetic variant associations with cognitive performance across development in British birth cohorts. Nat. Hum. Behav. 10, 1841&#x2013;1854 &#010;                https:\/\/doi.org\/10.1038\/s41562-026-02491-8&#010;                &#010;               (2026).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR82\" id=\"ref-link-section-d310663722e3724\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> by fitting a one-factor model and calculating scores using \u2018factanal\u2019 and Bartlett scoring in R using the following measures: Bracken School Readiness at age 3 years, reading vocabulary at ages 3 and 5 years, pattern construction at ages 5 and 7 years, and word reading and progress in math at age 7 years. The summarized measure explained 39% of the variance, and we included a total of 5,621 unrelated children of genetically inferred European ancestry with at least one genotyped parent. In INTERVAL, participants completed an FI test identical in nature to the one completed by UKB participants in the first in-person wave at two time points 12 months apart (test\u2013retest correlation\u2009=\u20090.65, P\u2009&lt;\u200910\u221215). We took the average of the FIS for individuals who had two measurements. We included 20,328 unrelated individuals of genetically inferred European ancestry. All variables were standardized to have a mean of 0 and variance of 1.<\/p>\n<p>Calculating PGIs<\/p>\n<p>We used LDpred2-auto<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 86\" title=\"Priv&#xE9;, F., Albi&#xF1;ana, C., Arbel, J., Pasaniuc, B. &amp; Vilhj&#xE1;lmsson, B. J. Inferring disease architecture and predictive ability with LDpred2-auto. Am. J. Hum. Genet. 110, 2042&#x2013;2055 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR86\" id=\"ref-link-section-d310663722e3741\" rel=\"nofollow noopener\" target=\"_blank\">86<\/a> to calculate PGIs in ALSPAC, MCS and INTERVAL using the GWAS for observed FIS (that is, measured average FIS), combined FIS and combined FIS\u2009+\u2009COGENT GWAS meta-analysis. We also used unrelated individuals with European ancestries (<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">Supplementary Methods<\/a>) and generated LD reference panels restricted to HapMap 3\u2009+\u2009SNPs<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 86\" title=\"Priv&#xE9;, F., Albi&#xF1;ana, C., Arbel, J., Pasaniuc, B. &amp; Vilhj&#xE1;lmsson, B. J. Inferring disease architecture and predictive ability with LDpred2-auto. Am. J. Hum. Genet. 110, 2042&#x2013;2055 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR86\" id=\"ref-link-section-d310663722e3748\" rel=\"nofollow noopener\" target=\"_blank\">86<\/a>, which are designed to maximize genome-wide tagging coverage, improving performance across diverse ancestries compared to the original set. We used default parameters to calculate SNP weights for the PGIs. We then imputed missing parental genotypes in both cohorts (that is, if one but not both parents were genotyped) using SNIPAR<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Young, A. I. et al. Mendelian imputation of parental genotypes improves estimates of direct genetic effects. Nat. Genet. 54, 897&#x2013;905 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR33\" id=\"ref-link-section-d310663722e3752\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a> as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 82\" title=\"Malawsky, D. S. et al. Common and rare genetic variant associations with cognitive performance across development in British birth cohorts. Nat. Hum. Behav. 10, 1841&#x2013;1854 &#010;                https:\/\/doi.org\/10.1038\/s41562-026-02491-8&#010;                &#010;               (2026).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR82\" id=\"ref-link-section-d310663722e3756\" rel=\"nofollow noopener\" target=\"_blank\">82<\/a> and calculated PGIs with the SNP weights generated above using the pgs.py script. We filtered PGI SNPs to those with INFO score &gt; 0.98 in each given cohort.<\/p>\n<p>Estimating direct and population effects of PGIs on cognitive performance measures<\/p>\n<p>To estimate the population effects of the PGIs, we standardized the PGIs as above and regressed the phenotypes on the individual\u2019s PGI, 20 genetic PCs and sex, and additionally age and age2 in INTERVAL, using the \u2018lm\u2019 function in R. To estimate the direct effects in ALSPAC and MCS, we added the two parental PGIs and additional covariates to the previous regression. The coefficient estimated for the child\u2019s PGI is the partial correlation between the PGI and the phenotype, and represents the direct genetic effect.<\/p>\n<p>In MCS, we accounted for ascertainment biases due to the cohort\u2019s nonrandom sampling scheme and attrition by incorporating sampling weights as previously described<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 8\" title=\"Huang, Q. Q. et al. Examining the role of common variants in rare neurodevelopmental conditions. Nature 636, 404&#x2013;411 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR8\" id=\"ref-link-section-d310663722e3773\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>. Briefly, we generated nonresponse weights using inverse probability weighting and multiplied these weights with the full UK sampling weights generated by the study. We then used these weights in the regression analyses conducted in MCS using the weights argument in the \u2018lm\u2019 function in R.<\/p>\n<p>Multilevel mixed-effects regression<\/p>\n<p>We used a multilevel mixed-effects regression to evaluate the association between discovery GWAS sample size and standardized PGI effect sizes using the \u2018metafor\u2019 package in R<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 87\" title=\"Viechtbauer, W. Conducting meta-analyses in R with the metafor package. J. Stat. Softw. 36, 1&#x2013;48 (2010).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR87\" id=\"ref-link-section-d310663722e3786\" rel=\"nofollow noopener\" target=\"_blank\">87<\/a>. We regressed the PGI \u03b2s, weighted by their inverse variance to account for measurement error, on the logarithm of the GWAS discovery sample size.<\/p>\n<p>Random intercepts were included for the cohorts (ALSPAC, MCS and INTERVAL) and for effect type (population and direct) nested within each cohort to account for heterogeneity in effect estimates across these groups. The model was specified as:<\/p>\n<p>$${\\beta }_{\\left(\\mathrm{ijk}\\right)}={\\gamma }_{0}+{\\gamma }_{1}\\mathrm{log}\\left({n}_{\\left(\\mathrm{ijk}\\right)}\\right)+{u}_{{\\rm{i}}}+{v}_{{\\rm{j}}}\\left(i\\right)+{\\varepsilon }_{\\left(\\mathrm{ijk}\\right)}$$<\/p>\n<p>where \u03b30 is the overall intercept, \u03b31 represents the fixed effect of log(n), u\u1d62 is the random effect for the ith cohort, v\u2c7c(i) is the random effect for effect type nested within the ith cohort and \u03b5\u208d\u1d62\u2c7c\u2096\u208e is the residual error. We also ran the model using direct effects estimates alone, dropping the v\u2c7c(i) term.<\/p>\n<p>Enrichment of de novo damaging variants in probands with neurodevelopmental conditions<\/p>\n<p>We used previously published data from ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 21\" title=\"Kaplanis, J. et al. Evidence for 28 genetic disorders discovered by combining healthcare and research data. Nature 586, 757&#x2013;762 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR21\" id=\"ref-link-section-d310663722e3936\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a> on three cohorts totaling over 31,000 exome-sequenced probands with developmental conditions and their parents to assess enrichment of damaging de novo mutations in the set of 12 genes with rare variant associations with FIS (FDR\u2009&lt;\u20091%) that are not DDG2P genes.<\/p>\n<p>We evaluated enrichment of de novo synonymous and nonsynonymous variants in the probands as follows. We used the null mutational model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Samocha, K. E. et al. A framework for the interpretation of de novo mutation in human disease. Nat. Genet. 46, 944&#x2013;950 (2014).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR36\" id=\"ref-link-section-d310663722e3943\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a> to compute, for each consequence class i, the cumulative per-site mutation rate \u03bci,g across all callable variants of that class in a given gene g, that is, the summed mutation rate across all sites in a gene. The expected burden of de novo mutations of class i for a given gene-set G was then calculated as:<\/p>\n<p>$${E}_{i,G}=\\sum _{g\\in G}{n}_{g}{\\mu }_{i,g},$$<\/p>\n<p>where ng is the total number of probands. We compared the observed count Oi of de novo mutations in class i to Ei by performing a one-sided Poisson test with rate parameter \u03bb\u2009=\u2009Ei, calculating P(X\u2009\u2265\u2009Oi|X\u2009~\u2009Pois(\u03bb)) as the significance of any excess de novo mutations. This framework allows assessment of whether damaging protein-truncating or missense variants occur more often than expected by chance, with synonymous variants serving as an internal negative control.<\/p>\n<p>Replication of gene-based burden test results in ALSPAC and MCS<\/p>\n<p>To replicate the results of our gene-based tests in UKB, we performed gene-set burden tests in ALSPAC and MCS. For both cohorts, we included only variants with a MAF lower than 0.1% in participants with European ancestry and had a gnomAD (v3; ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434&#x2013;443 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR18\" id=\"ref-link-section-d310663722e4097\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a>) allele frequency of\u2009&lt;\u20093\u2009\u00d7\u200910\u22125 (corresponding to allele count\u2009&lt;\u20095). As in UKB, we only retain high-confidence PTVs defined by LOFTEE<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 18\" title=\"Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434&#x2013;443 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR18\" id=\"ref-link-section-d310663722e4103\" rel=\"nofollow noopener\" target=\"_blank\">18<\/a> and missense variants with an MPC score &gt; 2. We then calculated the burden of PTVs for the following three gene sets: (1) all 26 genes discovered in UKB at FDR\u2009&lt;\u20091%; (2) the FDR\u2009&lt;\u20091% genes excluding the eight already identified in ref. <a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 20\" title=\"Chen, C.-Y. et al. The impact of rare protein coding genetic variation on adult cognitive function. Nat. Genet. 55, 927&#x2013;938 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#ref-CR20\" id=\"ref-link-section-d310663722e4107\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>; and (3) the 21 FDR\u2009&lt;\u20091% genes that we only discovered when using combined FIS. We used linear regression models to examine the association between cognitive measures and the gene-set burden of PTVs in unrelated children with European ancestry in both the Avon Longitudinal Study of Parents and Children (ALSPAC; n\u2009=\u20095,283) and the MCS (n\u2009=\u20095,621), controlling for sex and population structure with the top ten PCs.<\/p>\n<p>Reporting summary<\/p>\n<p>Further information on research design is available in the <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41588-026-02787-5#MOESM2\" rel=\"nofollow noopener\" target=\"_blank\">Nature Portfolio Reporting Summary<\/a> linked to this article.<\/p>\n","protected":false},"excerpt":{"rendered":"Ethics oversight UK Biobank obtained ethical approval from the NHS North West Centre for Research Ethics Committee (reference&hellip;\n","protected":false},"author":2,"featured_media":793025,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[10],"tags":[5083,5085,3251,5082,59,5084,3250,18005,102,5081,18006,13781,46180,56,54,55],"class_list":["post-793024","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health","tag-agriculture","tag-animal-genetics-and-genomics","tag-biomedicine","tag-cancer-research","tag-gb","tag-gene-function","tag-general","tag-genome-wide-association-studies","tag-health","tag-human-genetics","tag-neurodevelopmental-disorders","tag-population-genetics","tag-sequencing","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/793024","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=793024"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/793024\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/793025"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=793024"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=793024"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=793024"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}