Yala · Science translational medicine 2021 · Multi-center retrospective validation study · n=?

Toward robust mammography-based models for breast cancer risk.

Cited 277 times in the scientific literature.

Level 3 - non-randomized controlled study

Retrospective multi-center cohort validation study of a prognostic algorithm

PubMed 33504648 · doi:10.1126/scitranslmed.aba4373 · record verified 2026-08-28

What was done

Researchers developed Mirai, a mammography-based deep learning model designed to predict breast cancer risk across multiple time points, incorporate missing clinical risk factors, and remain consistent across mammography hardware. Mirai was trained on imaging data from Massachusetts General Hospital (MGH, USA) and evaluated on held-out test cohorts from MGH, Karolinska University Hospital (Sweden), and Chang Gung Memorial Hospital (CGMH, Taiwan). Performance was compared with the Tyrer-Cuzick clinical model and two earlier deep learning models (Hybrid DL and Image-Only DL).

What was found

Mirai achieved C-indices of 0.76 (95% CI, 0.74 to 0.80) at MGH, 0.81 (95% CI, 0.79 to 0.82) at Karolinska, and 0.79 (95% CI, 0.79 to 0.83) at CGMH. Its 5-year ROC AUC was significantly higher than Tyrer-Cuzick (P < 0.001), Hybrid DL (P < 0.001), and Image-Only DL (P < 0.001). On the MGH test set, Mirai classified 41.5% (95% CI, 34.4 to 48.5) of patients who developed breast cancer within 5 years as high risk, compared to 36.1% (95% CI, 29.1 to 42.9; P = 0.02) for Hybrid DL and 22.9% (95% CI, 15.9 to 29.6; P < 0.001) for the Tyrer-Cuzick model.

Why it matters

Mirai demonstrates that mammography-based deep learning risk models can generalize across distinct international health systems and hardware, substantially outperforming traditional clinical risk scores in identifying future cancer cases.

Limits

The abstract does not state the sample sizes or demographic characteristics of the training and validation cohorts. The analysis was retrospective, meaning prospective clinical utility, impact on screening workflows, and ultimate diagnostic outcomes remain to be established.

Cited by