Disease prediction with multi-omics and biomarkers empowers case-control genetic discoveries in the UK Biobank.
Level 3 - non-randomized controlled study
Non-randomized prospective cohort analysis leveraging longitudinal biobank data
PubMed 39261665 · doi:10.1038/s41588-024-01898-1
What was done
Researchers developed MILTON (machine learning with phenotype associations), an ensemble machine-learning framework utilizing multi-omics biomarkers to predict incident cases across 3,213 diseases from UK Biobank longitudinal health records. They applied MILTON to augment phenome-wide association studies across 484,230 genome-sequenced individuals (including 46,327 with matched plasma proteomics) and validated discoveries in the FinnGen biobank alongside two orthogonal machine-learning methods.
What was found
MILTON predicted incident disease cases undiagnosed at recruitment and was reported to outperform available polygenic risk scores. Incorporating the model into genetic association analyses improved signals for 88 known gene-disease relationships (P < 1 × 10^-8) and revealed 182 gene-disease relationships that did not reach genome-wide significance in non-augmented baseline cohorts.
Why it matters
Combining multi-omics biomarkers with machine learning enhances statistical power in genetic association studies, enabling the discovery of disease-associated genes that standard case-control analyses miss.
Limits
The abstract reports no numerical accuracy or discrimination metrics (such as AUC, sensitivity, or specificity) for the disease prediction models. In addition, the UK Biobank and FinnGen cohorts are predominantly of European ancestry, which may limit generalizability to more diverse populations.
Cited by
- supports AstraZeneca developed an AI model named Milton using UK Biobank plasma protein data to accurately predict cancer and neurodegenerative diseases up to 10 years before onset.