Ethical approval

This study involves data from human participants that have participated in the OFH program and the UKB. Ethical approval for OFH was granted by the Cambridge East Research Ethics Committee (REC reference: 21/EE/0016). The UKB has approval from the North West Multi-Centre Research Ethics Committee as a Research Tissue Bank. All participants provided written informed consent before participation in either study.

Study populations

OFH is a nonprobability, volunteer-based prospective cohort. The sampling frame comprises the UK-resident adult population (aged ≥18 years). Participants were recruited via in-person appointments at centers and pop-up clinic locations across England, Wales and Scotland and through invitation letters sent to all individuals and households, typically within a 5-mile radius of a newly opened center23 (Fig. 1c; recruitment in Northern Ireland only begun in late 2025), drawn from more than 2.5 million participants who have consented to take part (Extended Data Table 2). The recruitment approach first prioritized the inclusion of participants from more deprived areas and from UK minoritized ethnic groups that have been historically underrepresented in health research (Fig. 1), as recently described23. Our analyses included approximately 1.9 million participants (total sample size varies slightly across analyses owing to variation in samples sizes across data sources and differing exclusion criteria). Available data from consented participants included responses from the baseline questionnaires (participants completed either version 1 or version 2), which comprised up to 288 questions across five sections (Fig. 1b), along with clinical measurements, geolocation data and linked EHRs. OFH releases new data on a quarterly basis; we relied on OFH data from releases 12 and 13, released on 17 September and 11 December 2025, respectively. Information on the latest data release is available at https://ourfuturehealth.gitbook.io/our-future-health/data-releases/current-data-release.

In our study, analyses were restricted to participants with a valid age at registration (≥18 years). For analyses including ethnicity, participants who self-reported ‘do not know’, ‘prefer not to answer’ or did not provide a response were excluded. Analyses of questionnaire items unique to version 2 (‘Phenotyping’ section) were restricted to participants who completed either v2.1 or v2.2 of the questionnaire as part of data release 12 (97.2%; N = 1,847,750), unless otherwise indicated. Percentages are at times reported relative to the full cohort for consistency with cohort-wide reporting; therefore, totals may not always sum to 100%. Throughout, prevalence is defined as the proportion of OFH participants reporting or diagnosed with a particular condition, ascertained from the available data sources used in each analysis (for example, self-report or linked health records; ‘Phenotyping’ section). Unless otherwise specified, age refers to self-reported age at registration; sex refers to self-reported sex assigned at birth; and ethnicity refers to self-reported ethnicity using the list of Level 2 ethnic groups defined by the UK Government and adapted by OFH.

To situate the prevalence of health conditions and diseases in OFH within the context of other international cohorts, we also compared prevalence for multiple well-studied common diseases (N = 13) to three other national, large-scale biobanks for which we had access to comparable data: All of Us4, FinnGen6 and the UKB3. Given the UK-focus of OFH, and the fact that the baseline questionnaire was intended to align closely with the UKB (https://ourfuturehealth.org.uk/our-research-mission/how-our-future-health-works/), we conducted further cross-cohort comparisons between both cohorts. The UKB is a prospective, population-based study of half a million participants aged 40–69 years recruited across the UK between 2006 and 2010. The cohort design, recruitment procedures and data collection methods are described elsewhere3. Although the UKB has been found to not be suitable for deriving generalizable disease prevalence15, its large size and deep phenotyping have made it a valuable resource for public health research2; hence, we focus on it here. To assess comparability between cohorts, we performed sensitivity analyses matching the OFH and UKB cohorts on key demographic factors (restricted to participants aged 40–69 years who self-reported sex assigned at birth as male or female) (Extended Data Fig. 3 and Supplementary Tables). Notably, as UKB and OFH participants of equivalent age belong to different birth cohorts, several observed divergences in condition prevalence may still reflect birth cohort effects relating to generational shifts in exposure or disease risk.

To compare OFH with the general UK population, we selected nationally representative survey data that best matched the demographic composition (sex, age and country of registration) and initial recruitment period of OFH. When characteristics in national survey summaries were reported only in aggregated age, sex subgroups or only for specific countries of the UK, OFH participants were stratified or restricted to corresponding categories for comparability. Following ref. 23, our comparisons focus primarily on England, Wales and Scotland, where OFH recruitment invitation and in-person appointments started (appointments in England began in late 2022 and appointments in Wales and Scotland began in early 2024). Recruitment in Northern Ireland only begun in late 2025, and hence, data from participants in this country are not included in data releases 13 or 12 used in this study; the small percentage of participants who self-reported living in Northern Ireland (0.2%) at the time of registration probably represent participants who traveled to England, Scotland or Wales for a clinic measurements appointment.

Data on ethnicity were obtained from the 2021 census for England and Wales55 and the 2022 census for Scotland56,57, referred to throughout as the 2021–2022 UK census for simplicity. Anthropometric, lifestyle and health-related characteristics were obtained from the HSE for the years 2021 and 2022 (ref. 54). The HSE is an annual cross-sectional survey of a representative sample of private households in England, selected through a two-stage random probability design, which collects data on health conditions and health-related behaviors. Data on disease prevalence in the general UK population were obtained from the most recent GBD study at the time of writing58. The methodology for GBD estimates has been described in detail elsewhere59. Analyses were restricted to the year 2021. Data on national cancer prevalence were obtained from the NDRS60, which provides counts and rates of individuals living in England with or beyond a cancer diagnosis as of 31 December 2022 (for diagnoses since 1 January 1995).

Although not a focus herein, OFH also already provides a growing genetic and biological resource21, blood samples for genotyping on a custom 700,000-variant array, with approximately 750,000 participants with genotyped and imputed genotype array data. Other samples will in time enable additional omics assays (for example, proteomics, whole-genome sequencing (WGS) and methylation)23. OFH therefore already provides a complementary resource to existing resources through its combination of very large-scale active recruitment, data linkage and infrastructure for translational research. However, its relative utility compared with established large-scale cohorts such as All of Us, China Kadoorie Biobank, Million Veteran Program and UKB will still depend on the specific research question, study design and data modalities required (see Supplementary Table 13 for a brief overview of these biobanks compared with OFH).

Phenotyping

Owing to the timing of our study and only a marginal increase in baseline phenotypic data between data releases (1.5%), we relied on data release 12 for most primary analyses, which included: (1) self-reported data from participants who submitted participant demographic data and either version 1 (N = 52,745) or version 2 (N = 1,847,812) of the baseline questionnaire (version 2 includes additional questions and allows participants to respond to new or updated questions; further details and OFH’s versioning approach are described at https://ourfuturehealth.gitbook.io/our-future-health/data/questionnaire-data) on or before 15 July 2025 (total N = 1,900,557); (2) clinic measurements data from participants who attended an appointment on or before 1 April 2025 (N = 1,433,275); and (3) linked EHRs (including inpatient visits, outpatient visits, cancer registry and death registration data available for participants registered in England) for participants who completed a questionnaire before 9 April 2025 and were successfully linked to a NHS number (N = 1,703,250), with N = 1,665,668 having at least one linked secondary care or death registration record. These records span multiple EHR sources and time windows; therefore, case ascertainment and follow-up vary by dataset. We relied on data release 13 to analyze the geographic distribution of OFH participants (N = 1,841,458), as it included significantly more self-reported country and regional location data compared with data release 12 (N = 102,103) and the prevalence of selected uncommon and rare conditions (N = 187) in participants with linked EHRs (N = 1,690,845). To maximize sample size, all available data were included where possible; consequently, total sample size varied across analyses.

To characterize the OFH cohort, we used participant self-reported data on sociodemographic indicators, region of registration, lifestyle exposures, family health history and personal health history including past diagnoses and current medication use, obtained from the baseline questionnaire. Data on height, weight and waist circumference were obtained from the clinic measurements data. Exact details of the questionnaire and how the clinic measurements are taken are described elsewhere23. To provide examples of how disease end points can be constructed from different OFH datasets and explore prevalence of rare conditions, we further used data from EHRs, specifically, inpatient visits (admitted patient care episodes at NHS hospitals in England) and outpatient visits (outpatient appointments at NHS hospitals in England) collected by NHS England and linked to OFH. We also used linked cancer registry (patient tumor-level) data from the NDRS and linked death registration (cause of death) mortality data collected by the Office for National Statistics (ONS). Cancer prevalence was calculated in OFH for all cancers combined (excluding nonmelanoma skin cancer) and common types (breast, prostate, colon/rectal and lung or bronchial cancer) using self-report and EHRs, comprising both inpatient visits and linked cancer registry patient tumor data. Diagnostic information in linked EHR datasets was encoded using the ICD-10, and case ascertainment included primary and secondary diagnoses. For cancer analyses, we restricted case ascertainment to primary diagnoses recorded in linked cancer registry and inpatient datasets. Full details of all OFH data that has been measured and linked is available on the dedicated researcher website (https://research.ourfuturehealth.org.uk/data-and-cohort/).

We additionally derived several phenotypes using the Python library Phenofhy51, created as a companion to this paper and for other OFH users, including: age at registration, smoking status, BMI, physical walking activity (walking ≥16 times a month for at least 10 min) and multiple medication-use variables (number of medication categories used, polypharmacy, proportion of medication systems used and use across ≥2 systems). Where possible, we also sought to harmonize categorical variables to be consistent with UKB reporting15 for ease of comparison (see ‘Derived phenotypes’ section for details of the derivation logic).

To compare case prevalence for (N = 13) well-studied diseases between OFH, All of Us, FinnGen and the UKB, phenotypes were defined using ICD-10 codes in linked EHR. Data for FinnGen and the UKB were obtained from prior work6, and data for All of Us were obtained from the All of Us data snapshots. For cross-cohort prevalence comparisons, harmonized ICD-10 phenotype definitions were applied consistently across datasets where possible, with case definitions based on the presence of at least one qualifying ICD-10 diagnosis code. To robustly compare prevalence of diseases and health conditions between OFH and the UKB, cases were identified using participant self-reported diagnoses in both cohorts. For each self-reported disease or health condition in OFH, we reviewed each possible match in UKB non-cancer illness codes or ever diagnosed mental health conditions and linked it to the equivalent self-reported condition. A total of N = 109 shared phenotypes were identified. In the absence of standardized mappings (see ref. 21 for a discussion), to compare prevalence between OFH and the UK population using the GBD, we mapped self-reported health conditions to GBD level 3 and 4 causes that were clinically equivalent and rejected any mapping where aggregation would distort comparisons. This conservative approach resulted in a total of N = 45 shared phenotypes.

Finally, to provide an initial, exploratory assessment of case numbers for less common and rare conditions in OFH participants, we analyzed linked inpatient EHRs using the ICD-10-to-ORPHA code61 (an extensive online resource for rare diseases) mappings adapted from the consensus framework approach used in prior work examining rare diseases in the UKB. This involved using mappings in which the selected ORPHA code is as specific as possible, but no more specific than the corresponding ICD-10 code, thereby enabling reliable identification of rare conditions from ICD-10-coded diagnoses (see section 2.16 in the Supplementary Information for further details). In OFH, we analyzed N = 187 conditions previously characterized in UKB30 (Supplementary Table 16), including clinically defined rare and ultra-rare diseases as well as selected low-prevalence conditions retained for biobank comparison purposes. Cases were defined as participants with at least one qualifying ICD-10 code recorded in either primary or secondary inpatient diagnoses corresponding to the mapped ORPHA definitions. Prevalence was estimated by dividing the number of individuals with mapped ICD-10-coded diagnoses by the total number of OFH participants with linked inpatient EHR available. Further details of all mapping processes, including limitations, are discussed in the Supplementary Information, and the full list of mapped phenotypes are provided in the Supplementary Tables.

Derived phenotypes

Age at registration was derived from self-reported date of birth and date of registration. BMI was calculated as weight (kg) divided by height squared (m2), using clinic measurements (Supplementary Tables 4.1 and 4.2). Self-reported alcohol-consumption categories were harmonized to UKB conventions15 to support cross-cohort comparison (Table 1 and Supplementary Tables 4.1 and 4.2).

Smoking status was derived using questionnaire-specific logic. Current smokers reported cigarette smoking at any frequency. Never smokers reported no current smoking and fewer than 100 lifetime smoking occasions (including non-cigarette products). Former smokers reported no current smoking but either daily or most-days smoking previously or at least 100 lifetime occasions. Questionnaire-version-specific variables are listed in Supplementary Tables 4.1 and 4.2 footnotes.

Physical activity was defined as walking at least 16 times per month for at least 10 min, derived from questionnaire version 2. Medication-use indicators and aggregated medication-pattern variables were derived using the Phenofhy package and are summarized in Supplementary Tables 8, 9 and 12.

Standard exclusions

Participants with invalid age at registration (coded −999) were excluded from age-adjusted analyses (all regression models; Supplementary Tables 3 and 9). Sex-specific analyses excluded participants reporting intersex, ‘prefer not to answer’, or missing sex. Analyses involving ethnicity excluded ‘do not know’, ‘prefer not to answer’ and missing responses.

Variables introduced only in questionnaire version 2 were analyzed among questionnaire version 2 respondents (Supplementary Tables 4–9). For cohort-wide descriptive reporting (Table 1 and Supplementary Tables 4 and 5), several percentages are presented relative to the full cohort; version-restricted denominators are noted explicitly in table footnotes.

Alongside standard exclusions for quality control, we also applied additional exclusions to comply with the OFH Statistical Disclosure Control (SDC) framework, which aims to minimize the risk of participant re-identification. In line with the SDC data guidelines (https://dnanexus.gitbook.io/ofh/airlock/safe-outputs-data-guidelines-for-export), throughout our Results and Supplementary Tables reporting, we do not report extreme values (for example, minimum and maximum values) or exclude frequencies or counts relating to all individuals in the study population or all but one individual, and we replaced cell values with ‘<10’ for counts or frequencies based on a small group (<10 observations).

To further minimize the risk of participant re-identification risk (disclosure), all research using the OFH TRE additionally follows the four-eyes principle, meaning that every output was scrutinized by two people at a minimum, including at least one person not involved in the research.

Statistical analysis

We summarized the sociodemographic, lifestyle and health-related characteristics of OFH participants and compared these distributions with estimates from the nationally representative HSE 2021/2022. Continuous variables are reported as mean ± s.d. and categorical variables as counts and proportions; percentages were calculated relative to the number of participants with available data for each variable. Self-reported regional representation was compared with UK census estimates from 2021 to 2022 (Fig. 2a,b). Regional representation was visualized using a log2 representation index, defined as log2(OFH proportion/UK population proportion), displayed on a symmetric scale.

To assess correspondence in disease prevalence between OFH and the UKB, we compared the prevalence of shared self-reported diseases or health conditions (Fig. 2e, Extended Data Fig. 3 and Supplementary Table 6), as well as for rare conditions defined using the consensus ICD-10-to-ORPHA mappings we use (Extended Data Fig. 8 and Supplementary Table 16). Prevalence was plotted on a log–log scale after exclusion of missing values. The Pearson correlation coefficient (r) was calculated between log10-transformed prevalence estimates for OFH and UKB, and a log–log linear regression was fitted to characterize the overall multiplicative relationship. This analysis quantified the concordance of disease prevalence between OFH and UKB and highlighted deviations from the overall trend. We further compared disease prevalence in OFH with that in the UK general population using estimates from the GBD for the UK (Fig. 2f and Supplementary Table 7). The same log–log framework was applied, correlating log10-transformed frequencies between OFH and GBD using Pearson’s r and fitting a log–log linear regression to describe the relationship. This comparison assessed how closely cohort-level disease frequencies in OFH align with population-level estimates.

To evaluate the reproducibility of known associations between health conditions and clinical correlates29, we estimated adjusted ORs for selected correlates in both OFH and UKB using logistic regression models of the form: outcome = correlate + age + sex + ethnicity + income (Supplementary information section 2.3). Self-reported female sex, white ethnicity and household income <£18,000 served as reference categories, with ‘false’ as the reference for self-reported diagnoses. Responses of ‘do not know’ and ‘prefer not to answer’ were retained as levels, whereas participants with missing ethnicity or income were excluded, leaving a maximum sample size of N = 1,897,796 (Supplementary Tables 3.1 and 3.2 for counts across analyses). Family history variables for autism, bipolar disorder and schizophrenia were defined by reports of the condition in a biological parent or sibling in OFH and by reports of that condition in a first-degree relative in UKB. UKB and OFH ORs were plotted with 95% confidence intervals (CIs) and Pearson’s r and log–log linear regression were computed between log10-transformed ORs.

Ethnicity-stratified estimates were restricted to ethnic groups in which case counts exceeded 1% of the analytic sample, providing a prespecified threshold for statistical stability; where confidence intervals are wide, stratified findings should be interpreted with caution and regarded as hypothesis-generating. The large number of participants from diverse ethnic groups is a strength of this study, enabling the detection of interaction effects that would probably be missed in smaller studies. However, several ethnic groups remain underrepresented, and for certain outcome–exposure combinations, stratified analyses could not be performed or were restricted to a subset of groups, a limitation that should be considered when interpreting and generalizing these findings.

We next modeled the association between current self-reported medication use and corresponding self-reported diagnoses in OFH using multivariable logistic regression adjusted for age, sex, ethnicity and income (reference categories of female, white ethnicity, household income <£18,000; ‘false’ as the reference for diagnosis and medication self-report) (Fig. 3c). Observations with missing self-reported ethnicity, income and medication responses or diagnosis responses were excluded, leaving a maximum N = 1,867,370 participants (see Supplementary Table 9.1 for counts across analyses). Analyses where it was possible to harmonize questionnaire versions 1 and 2 are noted in the Supplementary Information. To contextualize potential misclassification in self-reported phenotypes, we additionally compared selected self-reported diagnoses with linked inpatient EHR-derived diagnoses for several common chronic conditions, including cancer, asthma, type 2 diabetes, depression, hypertension, high cholesterol and osteoarthritis, using an unweighted Cohen’s kappa with a 95% confidence interval62 (Extended Data Fig. 6).

Finally, to characterize interdependencies among self-reported medication-use indicators, we computed pairwise correlations among medication-domain flags and derived aggregated medication usage-pattern variables (Fig. 3d). Binary–binary, binary–continuous and continuous–continuous relationships were quantified using the φ (phi), point–biserial and Pearson correlation coefficients, respectively, computed over the intersection of nonmissing participants (Supplementary Table 13). We used self-reported medication usage instead of dispensed medicines in primary care, as the latter were not yet available at the time of our analyses, whereas the former were available for the whole sample. Analyses were descriptive and exploratory in nature; effect estimates, correlations and regression coefficients are reported with corresponding 95% confidence intervals or P values where relevant. We did not model clustering within households or relatedness between participants in the present analyses. Regression models were used to summarize associations rather than to make causal inferences; accordingly, estimates should be interpreted as participant-level summaries. Researchers producing descriptive summary statistics with OFH need to use sampling weights that adjust for the population of the variables they examine, which OFH has announced it is actively developing, benchmarked to the 2021–2022 UK census23. Although we did not attempt any exploratory reweighting, the potential of weights to substantially alter the magnitude or direction of associations can be expected to vary by phenotype63 and research question. As with the UKB, although certain groups are underrepresented, they are probably still sufficiently large to enable reliable associations and disease risk64. However, this should be evaluated once appropriate weights become available.

Further analytic details, including model specifications, variable definitions and variable coding names are provided in the Supplementary Information and Supplementary Tables. Figures were generated in Visual Studio Code and subsequently refined for presentation using BioRender (https://www.biorender.com) and Inkscape (version 1.4.3).

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.