Estimating the reproducibility of psychological science
Level 4 - case-series / case-control
Level 4 by design analogy; a cross-sectional multi-study empirical replication initiative evaluating 100 published psychology papers.
OpenAlex W1897139626 · doi:10.1126/science.aac4716
What was done
Replications of 100 experimental and correlational studies published across three psychology journals were conducted using high-powered designs and original materials when available. Replications were evaluated based on statistical significance rates, effect size comparisons, confidence interval overlap, and subjective ratings of replication success.
What was found
Replication effect sizes were half the magnitude of the original effects. Whereas 97% of original studies reported statistically significant results, only 36% of replications were statistically significant. Forty-seven percent of original effect sizes fell within the 95% confidence interval of the replication effect size, and 39% were subjectively rated as having replicated. Assuming no bias in original findings, meta-analytic combination of original and replication results yielded 68% statistically significant effects. Replication success was predicted better by the strength of original evidence than by original or replication team characteristics.
Why it matters
This project provides systematic, large-scale empirical evidence of substantial effect size attenuation and low replication rates in published psychological literature.
Limits
The sample was restricted to 100 studies from three specific psychology journals, limiting generalization to other disciplines or journals. Total participant counts, specific journal titles, and criteria for subjective replication assessments are not reported in the abstract.
Cited by
- supports Approximately 40% of scientific publications cannot be replicated.