Taeyoung Lee · JMIR mhealth and uhealth 2023 · prospective multicenter validation study · n=75

Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study

Cited 104 times in the scientific literature.

Level 3 - non-randomized controlled study

Prospective multicenter validation study comparing diagnostic accuracy of consumer devices against reference standard polysomnography

OpenAlex W4386899754 · doi:10.2196/50983 · record verified 2026-08-27

What was done

Researchers conducted a prospective multicenter validation study evaluating 11 consumer sleep trackers (CSTs) against concurrent in-lab polysomnography (PSG) across a tertiary hospital and a sleep clinic in Korea. Devices included 5 wearables (Google Pixel Watch, Galaxy Watch 5, Fitbit Sense 2, Apple Watch 8, Oura Ring 3), 3 nearables (Withings Sleep Tracking Mat, Google Nest Hub 2, Amazon Halo Rise), and 3 airable apps (SleepRoutine, SleepScore, Pillow). The 75 participants generated 3,890 hours of CST data and 543 hours of PSG recordings, totaling 349,114 analyzed epochs split across two device-group configurations to avoid cross-device interference.

What was found

Epoch-by-epoch sleep stage agreement varied considerably, with macro F1 scores spanning from 0.26 to 0.69 across devices. SleepRoutine performed best during wake and rapid eye movement (REM) stages, whereas the Google Pixel Watch and Fitbit Sense 2 performed best in deep sleep. Device categories exhibited distinct error profiles: wearables showed high proportional bias for sleep efficiency, while nearables showed high proportional bias for sleep latency. Subgroup analyses found macro F1 scores varied by BMI, baseline sleep efficiency, and apnea-hypopnea index, but differences between male and female participants were minimal.

Why it matters

This study directly benchmarks multiple classes of commercial sleep trackers against gold-standard polysomnography, revealing that accuracy is highly device- and stage-specific and influenced by user sleep characteristics.

Limits

The study evaluated 75 clinical participants from two centers in Korea, limiting generalizability to broader healthy populations in home environments. Devices were evaluated in laboratory conditions rather than free-living settings, and proprietary algorithms may change over time without public updates.

Cited by