Garg · Nature genetics 2024 · prospective cohort study with machine learning and genetic association analysis · n=484230

Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank.

Cited 109 times in the scientific literature.

Level 3 - non-randomized controlled study

Non-randomized prospective cohort analysis leveraging longitudinal biobank data

PubMed 39261665 · doi:10.1038/s41588-024-01898-1 · record verified 2026-08-29

What was done

Researchers developed MILTON (machine learning with phenotype associations), an ensemble machine-learning framework utilizing multi-omics biomarkers to predict incident cases across 3,213 diseases from UK Biobank longitudinal health records. They applied MILTON to augment phenome-wide association studies across 484,230 genome-sequenced individuals (including 46,327 with matched plasma proteomics) and validated discoveries in the FinnGen biobank alongside two orthogonal machine-learning methods.

What was found

MILTON predicted incident disease cases undiagnosed at recruitment and was reported to outperform available polygenic risk scores. Incorporating the model into genetic association analyses improved signals for 88 known gene-disease relationships (P < 1 × 10^-8) and revealed 182 gene-disease relationships that did not reach genome-wide significance in non-augmented baseline cohorts.

Why it matters

Combining multi-omics biomarkers with machine learning enhances statistical power in genetic association studies, enabling the discovery of disease-associated genes that standard case-control analyses miss.

Limits

The abstract reports no numerical accuracy or discrimination metrics (such as AUC, sensitivity, or specificity) for the disease prediction models. In addition, the UK Biobank and FinnGen cohorts are predominantly of European ancestry, which may limit generalizability to more diverse populations.

Cited by