Accuracy of Wristband Fitbit Models in Assessing Sleep: Systematic Review and Meta-Analysis
Level 1 - systematic review of randomized trials
Systematic review and meta-analysis of diagnostic validation studies against a reference standard (polysomnography)
OpenAlex W2980985183 · doi:10.2196/16273
What was done
A systematic review and meta-analysis adhering to PRISMA guidelines evaluated the accuracy of wristband Fitbit models in measuring sleep parameters and stages against gold-standard polysomnography (PSG). Searches covered CINAHL, Cochrane, Embase, MEDLINE, PubMed, PsycINFO, and Web of Science. Out of 3,085 candidate articles, 22 met inclusion criteria for the systematic review and 8 provided quantitative data for meta-analysis.
What was found
Nonsleep-staging (movement-only) Fitbit models significantly overestimated total sleep time (by ~7 to 67 minutes; effect size = -0.51, P < .001; I² = 8.8%) and sleep efficiency (by ~2% to 15%; effect size = -0.74, P < .001; I² = 24.0%), while underestimating wake after sleep onset (by ~6 to 44 minutes; effect size = 0.60, P < .001; I² = 0%) and showing poor specificity (0.10–0.52) despite high sensitivity (0.87–0.99). Newer sleep-staging models utilizing heart rate variability performed substantially better: they showed no statistically significant differences from PSG for total sleep time (P = .29), sleep efficiency (P = .19), or wake after sleep onset (P = .25), though they underestimated sleep onset latency (P = .03). Sleep-staging models achieved higher sensitivity (0.95–0.96) and specificity (0.58–0.69) for detecting sleep epochs.
Why it matters
Recent-generation consumer wearables combining movement and heart-rate tracking provide reasonable gross estimates of sleep parameters in free-living conditions. However, their modest specificity indicates they still frequently misclassify wakefulness as sleep and cannot substitute for clinical polysomnography.
Limits
Only 8 of the 22 included studies provided quantitative data suitable for meta-analysis. The abstract does not specify total participant sample size, participant health status (healthy vs. clinical sleep disorders), or environmental testing conditions (lab vs. home). Specificity in sleep-staging models remains suboptimal (0.58–0.69), limiting reliability for detecting wake periods.
Cited by
- supports Commercial wearable devices are inaccurate at measuring deep sleep and REM sleep stages.