This study was preregistered on the Open Science Framework (https://osf.io/2yznd) on 18 January 2024. Deviations from the preregistration are outlined in Supplementary Note 4. Individual cohorts followed a standard operating procedure (SOP) document (https://osf.io/jyk6d/). The overall study design is depicted in Fig. 4.

Ethics and inclusion

This study conforms to all relevant ethics regulations. This study was approved by the departmental ethics committee for the School of Psychological Sciences at Birkbeck, University of London (reference no. 2021007) on 27 October 2020. Each of the cohorts included in this research received ethical approval from their local ethics committees. Researchers from each of the cohorts included were involved in the study implementation and are included as authors on this paper.

Cohorts and measures

The sample sizes included in the GWAS for each sample entered into the meta-analyses are detailed in Supplementary Table 1.

MoBa

MoBa is a population-based pregnancy cohort study conducted by the Norwegian Institute of Public Health99. Participants were recruited from all over Norway from 1999 to 2008. The women consented to participation in 41% of the pregnancies. Blood samples were obtained from both parents during pregnancy and from mothers and children (umbilical cord) at birth100. The cohort includes approximately 114,500 children, 95,200 mothers and 75,200 fathers. The establishment of MoBa and initial data collection were based on a licence from the Norwegian Data Protection Agency and approval from the Regional Committees for Medical and Health Research Ethics. The MoBa cohort is currently regulated by the Norwegian Health Registry Act. The current study was approved by the Regional Committees for Medical and Health Research Ethics (2016/1702). We used the year of birth and child’s sex from the Medical Birth Registry, a national health registry containing information about all births in Norway from 1967 onwards101.

The quality-controlled genotype dataset from MoBa includes 76,577 children of European genetic ancestry and 6,981,748 SNPs that passed quality control102. We used phenotypic data from the 18- and 36-month data collection sweeps. After we filtered out individuals with no relevant phenotype data, there were n = 55,269–55,337 (27,010–27,042 females and 28,256–28,295 males) in the second-year analyses, n = 42,724–42,754 (20,903–20,913 females and 21,821–21,841 males) in the third-year analyses and n = 39,234–39,290 (19,218–19,250 females and 20,009–20,040 males) in the cross-age analyses.

ALSPAC

The ALSPAC (also known as ‘Children of the 90s’) study is a multi-generational observational prospective birth cohort29,103. Pregnant women resident in Avon, UK, with expected dates of delivery between 1 April 1991 and 31 December 1992 were invited to take part in the study. The initial number of pregnancies enrolled was 14,541, resulting in 13,988 children included in the study who were alive at 1 year of age. Of these pregnancies, some were from the same mother, meaning that there were 14,203 unique mothers who were initially enrolled in the study. Additional recruitment provided a total of 14,833 unique women enrolled in ALSPAC as of September 2021. Ethical approval for the study was obtained from the ALSPAC Ethics and Law Committee and the Local Research Ethics Committees. Consent for biological samples was collected in accordance with the Human Tissue Act (2004). Informed consent for the use of data collected via questionnaires and clinics was obtained from participants following the recommendations of the ALSPAC Ethics and Law Committee at the time.

After quality control and imputation of genetic data, 8,972 children (4,355 females and 4,572 males) remained in the sample, and 27,384,358 genetic variants were available for analysis. We used phenotypic questionnaire data from the 24- and 38-month data collection sweeps. The ALSPAC cohort was not included in the cross-age sociability analysis because the sociability phenotype was available in only the 38-month data collection sweep. After we filtered out individuals with no relevant phenotype data, there were n = 6,955–6,966 (3,382–2,287 females and 3,573–3,579 males) in the second-year analyses, n = 6,850–6,855 (3,341–3,343 females and 3,508–3,513 males) in the third-year analyses and n = 6,418–6,425 (3,127–3,131 females and 3,291–3,296 males) in the cross-age analyses.

The study website contains details of all the data that are available through a fully searchable data dictionary and variable search tool: http://www.bristol.ac.uk/alspac/researchers/our-data/.

TEDS

TEDS104 is a large community-based study of twins born in England and Wales that aims to investigate the relative genetic and environmental influences on individual differences in developmental, cognitive and emotional factors relevant to typical development. All twins born in England and Wales were invited to participate in TEDS through their parents; the twins were identified through the Office for National Statistics. The first wave of data collection took place when the twins were 18 months old, in which 13,759 families participated. After the first wave, data were collected at regular intervals throughout development and into adulthood. Informed consent was obtained through parents105. Analyses of TEDS data were conducted on the King’s College London CREATE HPC106.

Genetic data from TEDS twins were genotyped on two different genotype chips: Affymetrix GeneChip v.6.0 (n = 3,057 individuals) and Illumina HumanOmniExpressExome (n = 7,289 individuals). After quality control and imputation, 7,363,646 SNPs remained. Only one of each monozygotic twin pair was genotyped and is therefore included in these analyses. This analysis used data from the TEDS data collection waves at ages 2 and 3. After we filtered out those with no relevant phenotype data, there were n = 5,460–5,527 (2,825–2,861 females and 2,635–2,671 males) in the second-year analyses, n = 5,648–5,686 (2,907–2,924 females and 2,741–2,762 males) in the third-year analyses and n = 4,673–4,769 (2,422–2,472 females and 2,251–2,297 males) in the cross-age analyses.

The TEDS cohort was not included in the cross-age shyness analysis because the shyness phenotype was available in only the age 3 data collection sweep.

Lifelines

The Lifelines cohort study is a multidisciplinary prospective population-based cohort study examining in a unique three-generation design the health and health-related behaviours of 167,729 persons living in the North of the Netherlands. It employs a broad range of investigative procedures in assessing the biomedical, socio-demographic, behavioural, physical and psychological factors that contribute to the health and disease of the general population, with a special focus on multi-morbidity and complex genetics

Lifelines recruited individuals aged 25–50 residing in the North of the Netherlands between 2006 and 2013107. After their recruitment, in their initial study visit, individuals were asked for their consent for the study team to contact their family members with an invitation to participate. This method of recruitment thus gave rise to the three-generation design that the study employs, with cohort members as young as six months old.

Participants in the Lifelines study were sent a questionnaire every 1.5 years, which included a parent-report questionnaire for the children in Lifelines, in Dutch107. Only the activity phenotype was available for Lifelines. Questionnaires about the offspring in the sample were completed retrospectively by parents or caregivers, and questions specified the relevant time period in the child’s life on which to retrospectively report. After genetic data quality control and after we filtered out anyone without the relevant phenotype data, the sample size for Lifelines was n = 3,435 (1,780 females and 1,655 males).

The Lifelines cohort was included only in the analyses for the second year because it did not have relevant data for the third year.

NTR

NTR was established and started recruiting participants in 1986. There are two branches to this study, consisting of twins registered shortly after birth (Young NTR) and those recruited as adolescents or young adults30 (Adult NTR). The present study uses data from Young NTR108. In Young NTR, recruitment took place shortly after birth during a home visit to parents of newborn twins or multiples30. The parents completed questionnaires about their offspring in roughly two-year intervals from birth. For a selection of these twins, DNA samples were collected109. The genotyped sample in NTR consisted of n = 37,823 subjects. After quality control and the removal of anyone without relevant temperament phenotype data, the GWAS sample size was n = 1,433 (722 females and 711 males). We used phenotypic data from the data collection sweep when the twins were approximately two years old. The NTR cohort was included only in the analyses for the second year because it did not have relevant data for the third year.

TMM BirThree

The TMM study was established in 2012 in Japan110 to monitor the health impacts of the Great East Japan Earthquake and the tsunami that followed in 2011. Part of this project is TMM BirThree, a longitudinal birth cohort study that started recruitment in 2013111. TMM BirThree is a three-generational study: pregnant women, their partners (the father of the proband), their parents (the grandparents), any siblings of the proband and extended family members were recruited to the study. Participants were recruited between 2013 and 2017 when their delivery was scheduled (with an expected delivery date after 1 February 2014); they were recruited from the Miyagi prefecture, where the Great East Japan Earthquake happened, by their obstetrics clinics and hospitals. A total of 23,406 pregnancies were recruited, with some people participating with multiple subsequent births111. Exclusion criteria included non-consent, being unable to complete the study questionnaires and participation in another birth cohort study for that specific pregnancy. Questionnaires pertaining to the child were mailed to the parents; questionnaire data from the data collection sweep when the children were 24 months old were used in the present study.

The final sample included n = 5,618–5,653 individuals with relevant phenotypic data for this study (2,750–2,762 females and 2,868–2,891 males). These individuals were all genotyped on the Affymetrix Axiom Japonica v.2 array, which was designed for this specific population112.

DCHS

The DCHS is a population-based cohort study based in Drakenstein, South Africa, which is a peri-urban area outside of Cape Town113. Pregnant women were recruited during their second trimester from two primary care clinics in Drakenstein114. Enrolled mother–child dyads are being continuously followed up114. The study aims to investigate the early determinants of health, including neurodevelopment, in the context of low-income and middle-income countries, which are underrepresented and therefore underserved by research of this kind115. Phenotypic data from the DCHS used in the present study were collected during the data collection sweeps at 24 and 42 months.

This cohort included n = 595–843 genotyped individuals with relevant phenotype data: n = 595–596 (307 females and 288–289 males) for the second-year analyses and n = 843 (428 females and 415 males) for the third-year analyses. The genetic data (n = 956) were genotyped in three batches: two used the Infinium Global Screening Array (with one batch for cord blood, n = 152, and one for blood, n = 686), and one used the Infinium PsychArray (using cord blood, n = 118).

Phenotype measurement and coding

We preregistered to study four parent-rated infant and toddler temperament traits: emotionality, activity, shyness and sociability. Three cohorts, NTR, MoBa and ALSPAC (for the third-year data collection), included the EAS measure, and as such these scales were used as included in the cohorts. We note that NTR and ALSPAC used the scales as published, and MoBa’s EAS scales were reduced in items (see the instrument documentation at https://www.fhi.no/en/ch/studies/moba/for-forskere-artikler/questionnaires-from-moba/). For Lifelines, TEDS and ALSPAC (in the second year), we used appropriate items and scales from other parent-report temperament questionnaires that captured the same four EAS temperament constructs. ALSPAC (second-year data collection) used published subscales from other measures that captured the emotionality and activity constructs, and we therefore used these published subscales. Where cohorts did not include emotionality, activity, shyness or sociability subscales (that is, Lifelines, TEDS and the second year of ALSPAC for shyness), items that matched the EAS items were used to create subscales. See Supplementary Note 5 for further details of phenotype measurement in each cohort.

In all cohorts, the phenotypes were measured using Likert-scale questionnaires completed by parents or caregivers; for the details of questionnaire raters, see Supplementary Table 22. The individual’s phenotype score was calculated as a prorated sum score or mean of questionnaire items, scaled about a mean of 0 and standard deviation of 1. A completion rate of 50% was required; otherwise, the phenotype was coded as missing. In the cross-age analyses, the phenotype was calculated as a mean of the phenotype values from the second and third years. For phenotypic distributions of the cross-age phenotypes, see Supplementary Figs. 1921.

We conducted separate phenotypic analyses to verify the comparability of the different measures for the EAS traits where the EAS temperament survey was not employed by the cohorts. This did not include the phenotype measures from TMM BirThree and the DCHS (the Child Behaviour Checklist questionnaire). Briefly, in a separate infant and toddler sample (n = 118 for the second year; n = 84 for the third year), we collected parent-report data using all of the temperament measures used by the cohorts employed in this meta-analysis to assess concordance between phenotype measurements in different samples. The methods and results of these phenotypic analyses are provided in Supplementary Note 6.

In light of past research showing that genetic associations are stable across infancy and toddlerhood25,116, we have investigated temperament traits here both at specific ages (the second and third postnatal years) and as an average of the two years. Three cohorts had data for both ages: MoBa, ALSPAC and TEDS. For each cohort, the phenotype needed to be measured at both time points to be included in the cross-age analyses, because a mean of the phenotype across ages was used in these analyses. Sociability in ALSPAC and shyness in TEDS were thus not included in the cross-age analyses. The cross-age analyses were not preregistered.

Rater information

The MoBa and NTR phenotypes were measured using maternal-report questionnaires. ALSPAC used caregiver-report questionnaires; that is, the questionnaires could be completed by any combination of mother, father and/or ‘other’. The Lifelines phenotype was reported retrospectively by either biological parent or ‘other’. There was no evidence of significant differences between the mean reported phenotypes by different raters in the ALSPAC or Lifelines samples (Supplementary Note 7).

Quality control of genetic data

The main details of the procedures for individual cohort dataset genotyping, imputation and quality control that were conducted by the study teams are outlined in Supplementary Table 23 and are also detailed elsewhere29,102,105,107,109. An overview of the genotyping, ancestry filtering and imputation of the datasets as well as the quality-control procedures conducted for the present study can be found in Supplementary Table 23. Each sample underwent post-imputation quality control for the present study according to the guidelines from the Rapid Imputation for Consortias Pipeline (RICOPILI)117. Briefly, genetic variants with a low imputation score, minor allele frequency or call rate were removed, as well as variants that had a low P value for the Hardy–Weinberg equilibrium exact test. Individuals were removed if they had a low call rate, excess autosomal heterozygosity, an XXY genotype or other aneuploidies or if they fell outside of the relevant ancestry cluster when compared with a reference dataset (the filter thresholds are listed in Supplementary Table 23). Where X-chromosome data were available, reported sex was compared with sex according to X-chromosome heterozygosity, and individuals with mismatches were removed. The genetic relatedness of individuals was compared to the reported family structure, and any mismatches were removed as they were likely to be sample mix-ups.

GWAS

GCTA118 fastGWA119 was used to carry out genome-wide association analyses in each individual European-ancestry sample. In all association analyses, the first ten principal components and age (or date of birth as a proxy in MoBa) were included in the model as quantitative covariates, and sex and any relevant technical variables (for example, genotyping batch) were included as discrete covariates. A genetic relatedness matrix was also included to correct for relatedness structure within the sample. See Supplementary Note 8 for a full list of covariates included for each sample.

In each European-ancestry sample, GWAS analyses were completed for the whole sample, and also stratified by sex, for each of the four EAS traits in the second and third years and as a cross-age average.

GWAS meta-analysis

All summary statistics from individual sample GWAS went through quality control using GWASinspector (v. 1.6.5)120, which removed variants that had missing or invalid data in key columns (chromosome, position, β, β s.e., A1 and A2); were duplicated; were monomorphic, allosomal or mitochondrial; or had a low imputed information score (<0.8) where available. GWASinspector also compared variants to the 1000 Genomes reference panel121, removed variants that were not in this reference dataset and aligned variants to the reference strand.

The meta-analysis of summary statistics was conducted using METAL (version release date: 2020-05-05)80 using a standard-error-weighted fixed-effects model. Variants were included only if they had combined minor allele frequency >1% and if the rsID had matching positions between cohorts (if not, the sample variant from the sample with the higher n was retained). Meta-analyses were conducted for the whole-sample GWAS and for the sex-stratified GWAS. The genetic correlations between each of the pairs of sex-stratified meta-analyses for each phenotype were calculated. The total meta-analysed sample sizes are provided in Supplementary Table 1.

The results from all meta-analyses were filtered by the per-variant sample size. A filter was applied such that variants were required to have, at the least, an n totalling either the summed n of the cohorts not including MoBa, or 10,000, whichever was lower. The analyses for which the summed sample sizes of the cohorts not including MoBa were less than 10,000 were shyness in the second year, and shyness and sociability in the cross-age analyses.

Heterogeneity in the meta-analyses was assessed using METAL80, using I2 statistics and corresponding P values from Cochran’s Q-tests for each SNP. P values were assessed according to a genome-wide significance threshold of P < 5 × 10−8.

Multi-ancestry GWAS meta-analysis

We conducted GWAS meta-analyses including diverse-ancestry cohorts: TMM BirThree, a Japanese cohort study, and the DCHS, a cohort based in South Africa. The GWAS in TMM BirThree were conducted using GCTA fastGWA to account for known relatedness within the sample (see the methods above), and the GWAS in the DCHS was conducted using PLINK 2.0122 because there were no related individuals in the sample. These individual GWAS analyses and the meta-analyses including all cohorts used the same methods as those using only European-ancestry samples. The summary data underwent quality control using GWASinspector120 as above, and then the summary data were meta-analysed using a fixed-effects standard-error-weighted meta-analysis in METAL80. METAL was also used to conduct heterogeneity analyses, as above. SNPs were retained in the meta-analysed summary statistics if they had a minimum sample size of n = 10,000.

Details of the quality control conducted for these two diverse-ancestry samples are provided in Supplementary Table 23 as well as elsewhere115,123. Quality control of genetic data was conducted in line with the study SOP, as with the European-ancestry samples.

The phenotype in both cohorts was measured using the Child Behaviour Checklist parent-report questionnaire79. Prorated sum scores were calculated as above for the European-ancestry cohorts. See Supplementary Note 5 for further details.

Most post-GWAS analyses were not conducted using summary statistics from diverse-ancestry GWAS meta-analyses due to differences in LD structure between different ancestry populations and the need for accurate LD reference panels95.

Conditional analysis of lead variants

For GWAS that identified multiple genome-wide significant signals (P < 5 × 10−8) in the same locus, conditional analyses were performed using GCTA118 COJO51. This was done to identify whether there are multiple independent signals in a locus by conditioning on the lead SNP and testing for further signal independent of that SNP at the locus. Stepwise selection was used to identify independent lead variants in a single locus, with MoBa data used to estimate LD between SNPs.

FUMA

To further investigate the COJO independent SNPs as identified above, we carried out fine-mapping, functional annotation, chromatin-interaction mapping and gene-based analyses in FUMA53 (v.1.5.2) and MAGMA53,54 (v.1.08). A list of the COJO independent SNPs was input as a predefined set of lead SNPs. To carry out analyses on prioritized genes mapped to SNPs significantly associated with our phenotypes of interest, SNPs were defined to be independent if they had pairwise LD r2 < 0.6. LD blocks were merged into one locus if they were closer than 250 kb to one another. The 1000 Genomes Phase 3 EUR reference panel was employed by FUMA in these analyses.

Annotation was carried out using ANNOVAR within FUMA (downloaded on 17 July 2017). SNPs were mapped to genes at a maximum distance of 10 kb away using tissue types related to brain, nervous system, immune response, hormones, physical fitness and development for chromatin state mapping and annotation datasets using the following relevant HiC datasets: adult cortex and fetal cortex, adrenal, dorsolateral prefrontal cortex, hippocampus, lung, ovary, psoas, right ventricle, small bowel, spleen, mesenchymal stem cell, mesoderm and neural progenitor cell. We annotated enhancer/promoter regions from Roadmap 111 epigenomes of the following tissues: adrenal, brain, embryonic stem cells, embryonic stem cell derived, fat, column, duodenum, oesophagus, intestine, rectum, stomach, heart, lung, muscle, muscle leg, ovary, skin, spleen, stromal connective, thymus and induced pluripotent stem cells. Within FUMA, the GWAS Catalog124 (v.1.0.2, e112 r2024-06-17) was searched to identify prior GWAS for which candidate SNPs in the present GWAS were also significant.

eQTL mapping was performed in FUMA using the following relevant GTEx v.8 tissues: Adipose Subcutaneous, Adipose Visceral Omentum, Brain Amygdala, Brain Anterior cingulate cortex BA24, Brain Caudate basal ganglia, Brain Cerebellar Hemisphere, Brain Cerebellum, Brain Cortex, Brain Frontal Cortex BA9, Brain Hippocampus, Brain Hypothalamus, Brain Nucleus accumbens basal ganglia, Brain Putamen basal ganglia, Brain Spinal cord cervical_c-1, Brain Substantia nigra, Colon Sigmoid, Colon Transverse, Esophagus Gastroesophageal Junction, Esophagus Mucosa, Esophagus Muscolaris, Heart Atrial Appendage, Heart Left Ventricle, Lung, Muscle Skeletal, Nerve Tibial, Ovary, Pituitary, Cells Cultured Fibroblasts, Skin Not Exposed Suprapubic, Skin SunExposed Lower Leg, Small Intestine Terminal Ileum, Spleen, Stomach, Testis and Thyroid.

MAGMA gene-property analyses (one-sided test of a positive association between tissue specificity and genetic association with the temperament trait) were performed using GTEx v.8 (ref. 125) and BrainSpan126 (consisting of RNAseq data from 524 brain samples from 11 general developmental stages ranging from early prenatal to middle adulthood) tissue types. For MAGMA gene-set analysis, 4,761 curated gene sets and 5,917 Gene Ontology terms from MsigDB v.6.2 (ref. 127) provided within FUMA were used. The MHC region was excluded for all analyses except for activity in the second year and activity cross-age analyses, for which there were significant COJO independent SNPs in this region. SNPs were assigned to genes if they were within 1 Mb upstream and downstream53.

LDSC

SNP heritability and bivariate genetic correlations were calculated using LDSC72,74, using the 1000 Genomes Phase 3 (ref. 121) European genetic ancestry LD score reference panel. During data munging, variants were retained only if they matched a reference panel of high-quality Haplotype Reference Consortium variants. Two-sided tests were used to evaluate the null hypothesis (rg = 0).

Bivariate genetic correlations were calculated between each of the four EAS traits and other physical health, cognitive, neurodevelopmental, psychiatric and personality phenotypes. The traits included in these genetic correlation calculations were age at onset of walking46, BMI128, educational attainment58, cognitive performance (“intelligence”)67, lifetime anxiety disorder129, depression130, schizophrenia131, bipolar disorder (types 1 and 2, not including self-diagnosis)132, autism (pre-publication access to early freeze of upcoming GWAS), autism diagnosed before age 11 and autism diagnosed after age 10 (ref. 133), ADHD134 and the ‘big five’ personality traits, neuroticism, extraversion, agreeableness, conscientiousness and openness to experience135.

A Benjamini–Hochberg procedure was applied to control the FDR at 0.05. Any genetic correlations that reached significance under this criterion were followed up with MiXeR analyses (below; where further criteria were met).

MiXeR

Univariate causal mixture models were conducted using MiXeR76 to determine polygenicity estimates—that is, the number of SNPs that contribute to 90% of the SNP heritability. Bivariate causal mixture models were also fitted using MiXeR75 using all phenotype pairs identified by LDSC genetic correlation analyses. This was applied only when both phenotypes in the pair satisfied the criterion of h2SNP × n > 6,000 as specified by ref. 136 to ensure there was enough statistical power for the analyses. For all phenotypes in these analyses, we removed the MHC region (chr. 6: 26000000–34000000; GRCh37) due to its complex LD structure, and for the GWAS summary statistics for which a case–control design was employed, we calculated the effective sample size as neff = 4/(1/ncases + 1/ncontrols), as indicated by the authors. We used the authors’ summary statistics preparation pipeline to prepare all of the files being entered into both univariate and bivariate analyses (https://github.com/precimed/python_convert/).

Akaike information criterion values were used to compare the MiXeR model to a model that contains the minimum amount of polygenic overlap to explain the LDSC genetic correlation (‘minimal model’) for all bivariate MiXeR models. Positive differential Akaike information criterion values indicated that the MiXeR model was supported over the minimal model.

PGS

The variance explained by the PGS for each trait was calculated using a leave-one-out design. GWAS meta-analyses for each trait were conducted leaving out one sample in turn each time, and summary statistics were used to calculate PGS weights. The left-out sample was then used as a target sample to evaluate the PGS. This was repeated, leaving out each of the smaller samples (ALSPAC, TEDS, NTR and Lifelines), for each of the phenotypes separately. Given that the MoBa sample comprised the largest portion of the overall sample size, it would not be appropriate to leave it out as a target sample. A within-sample cross-validation method was therefore used for MoBa whereby one fifth of the sample was left out as a target sample and the rest of the sample used for a GWAS from which the summary statistics were employed to estimate the SNP weights.

All PGS were calculated using PRS-cs137,138, with weights derived per chromosome using the 1000 Genomes Phase 3 European panel121 as an LD reference panel. PLINK 2.0 (ref. 122) was used for computing PGS in the left-out target sample. The following PRS-cs parameters were used: the number of burn-in iterations was 500; the total number of MCMC iterations was 1,000; parameters a and b in the gamma-gamma prior were 1 and 0.5, respectively; the thinning factor of the Markov chain was 5; the global shrinkage parameter phi was 0.01; and the SNP effect sizes were standardized.

We quantified the proportion of variance explained using the adjusted R2 of the linear model regressing the PGS on the phenotype (both scaled about a mean of 0 and an s.d. of 1), controlling for ten genetic principal components and the genotyping batch (if applicable to the target sample)

Within- and between-family PGS analysis

We conducted within- and between-family PGS analyses to understand their relative contributions to each phenotype78. This was conducted in the TEDS and NTR samples (NTR only for specific-age analyses; the results are shown in Supplementary Fig. 14 and Supplementary Table 18), which contain DZ twin pairs, using a meta-analysis of the other samples to calculate PGS weights (as described above). We used data for DZ pairs with complete phenotypic data. First, bootstrapped intra-class correlations were calculated to ensure that the method was justified. The intra-class correlation represents the portion of the overall variance in the phenotype that can be accounted for by the within-pair similarity. When it tends to 0, the total variation is mostly attributed to individual-level variation. We used the method from Selzam et al.78, with scripts from Demange et al.139 used (https://github.com/PerlineDemange/GeneticNurtureNonCog/). For each individual, the between-family PGS was calculated as the mean PGS of that individual and their sibling pair. The within-family PGS was calculated as that individual’s PGS minus their between-family PGS.

We fitted a random-intercept mixed-effects linear model in R140, including age, sex, the first ten principal components and a genotyping batch dummy variable as covariates. A chi-squared test was used to compare the within- and between-family regression coefficients.

Colocalization

We used the coloc method64 with the default priors to assess whether GWAS and eQTL signals in the same region are likely to share a single causal variant. Colocalization was performed for each GWAS with at least one genome-wide significant SNP. For each significant locus, we tested all genes within ±1 Mb of the lead SNP, as well as genes prioritized by FUMA. eQTL data were derived from 1,433 post-mortem bulk RNA-seq samples of the adult human cortex65. A PP4 > 0.8 threshold was applied to identify significant colocalization events, where PP4 represents the posterior probability that both the GWAS and eQTL signals share a common causal variant. A PP4 value greater than 0.8 indicates strong evidence for colocalization, suggesting that the signals from both datasets are likely to be driven by the same underlying variant. Colocalization analyses were not performed for hits in the human leukocyte antigen (HLA) region on chromosome 6 due to the high LD complexity in this region, which complicates the identification of independent causal variants.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.