Lee · JMIR mHealth and uHealth 2023 · prospective multicenter validation study · n=75

Accuracy of 11 Wearable, Nearable, and Airable Consumer Sleep Trackers: Prospective Multicenter Validation Study.

Cited 104 times in the scientific literature.

Level 3 - non-randomized controlled study

Prospective multicenter diagnostic validation study comparing index tests to reference standard polysomnography.

PubMed 37917155 · doi:10.2196/50983 · record verified 2026-08-29

What was done

Seventy-five participants were recruited across two centers in Korea (a tertiary hospital and a specialized sleep clinic) to compare 11 commercial consumer sleep trackers (CSTs) against in-laboratory polysomnography (PSG). Devices included 5 wearables (Google Pixel Watch, Galaxy Watch 5, Fitbit Sense 2, Apple Watch 8, Oura Ring 3), 3 nearables (Withings Sleep Tracking Mat, Google Nest Hub 2, Amazon Halo Rise), and 3 smartphone airables (SleepRoutine, SleepScore, Pillow). Trackers were divided into two testing configurations (8 CSTs per group) to avoid signal interference. In total, 543 hours of PSG, 3,890 hours of CST recordings (averaging 353 hours per device), and 349,114 epochs were analyzed for epoch-by-epoch agreement and sleep metrics.

What was found

Epoch-by-epoch sleep stage classification performance varied widely across trackers, with macro F1 scores spanning from 0.26 to 0.69. Accuracy differed by sleep stage: SleepRoutine demonstrated superior performance in wake and REM stages, whereas wearables such as the Google Pixel Watch and Fitbit Sense 2 performed best in deep sleep. Systematic estimation biases emerged by device class: wearables showed high proportional bias for sleep efficiency, while nearables showed high proportional bias for sleep latency. Subgroup analyses showed macro F1 score variations by BMI, sleep efficiency, and apnea-hypopnea index, whereas differences between male and female participants were minimal.

Why it matters

This study provides head-to-head validation data across 11 modern wearable, nearable, and contactless sleep tracking technologies against the clinical gold standard. It demonstrates that consumer tracker accuracy depends strongly on the specific sleep stage, device modality, and user clinical characteristics such as sleep apnea.

Limits

Evaluations were conducted in a controlled hospital and sleep-clinic environment, which may not mirror unsupervised home tracking conditions. The cohort was limited to 75 participants in Korea presenting to clinical centers, and proprietary algorithm updates by manufacturers may alter device performance over time.

Cited by