Yland · Human reproduction (Oxford, England) 2022 · Prospective cohort study · n=4133

Predictive models of pregnancy based on data from a preconception cohort study.

Cited 24 times in the scientific literature.

Level 3 - non-randomized controlled study

Prospective cohort study developing and internally validating clinical prediction models

PubMed 35024824 · doi:10.1093/humrep/deab280 · record verified 2026-08-27

What was done

Researchers analyzed data from 4,133 female participants aged 21–45 years in the USA and Canada from a prospective preconception cohort (2013–2019) who were actively trying to conceive, had at most one cycle of attempt at entry, and were not using fertility treatments. Baseline questionnaires assessed 163 predictor variables spanning sociodemographics, lifestyle, diet, medical history, and partner characteristics, followed by bi-monthly questionnaires for up to 12 months or until conception. Machine learning algorithms (regularized logistic regression, support vector machines, neural networks, gradient boosted decision trees) and Cox models were used to predict pregnancy within <12 cycles (Model I), within 6 cycles (Model II), and per-cycle across 12 cycles (Model III).

What was found

Parsimonious models attained an AUC of 70% for Model I (<12 cycles), an AUC of 66% for Model II (6 cycles), and a concordance index of 63% for Model III (per-cycle). Consistent positive predictors across all models were previous infant breastfeeding and multivitamin or folic acid supplementation. Consistent inverse predictors were female age, female BMI, and history of infertility. For nulligravid women without past infertility, top predictors were female age, female BMI, male BMI, fertility app usage, attempt duration at study entry, and perceived stress.

Why it matters

This study shows that applying machine learning to broad preconception epidemiological data achieves modest discriminative performance (~70% AUC) for predicting natural conception, improving upon previous predictive models evaluated in subfertile cohorts.

Limits

Predictor variables were self-reported, creating potential for misclassification. Key unmeasured predictors may be missing. Validation relied on split-sample replication without external validation in an independent cohort.

Cited by