Associations between acoustic features of maternal speech and infants' emotion regulation following a social stressor.
Level 4 - case-series / case-control
Cross-sectional observational laboratory study (by design analogy, non-clinical developmental science)
PubMed 34618391 · doi:10.1111/infa.12440
What was done
Researchers assessed 94 mother-infant dyads (infant age 4–8 months) following the Still-Face paradigm social stressor. Infant physiological responses (heart rate and respiratory sinus arrhythmia/cardiac vagal tone) were derived from electrocardiogram (ECG) recordings, and behavioral distress was coded from negative vocalizations, facial expressions, and gaze aversion. Maternal vocalizations were analyzed using spectral analysis and spectro-temporal modulation via two-dimensional fast Fourier transformation of audio spectrograms to create a maternal prosody composite score.
What was found
High values on the maternal prosody composite were associated with: - Decreases in infant heart rate (β = -.26, 95% CI: [-0.46, -0.05]) - Decreases in infant behavioral distress (β = -.23, 95% CI: [-0.42, -0.03]) - Increases in cardiac vagal tone among infants with low vagal tone during the stressor (1 SD below mean β = .39, 95% CI: [0.06, 0.73]) Additionally, high infant heart rate predicted increases in the maternal prosody composite (β = .18, 95% CI: [0.03, 0.33]).
Why it matters
The findings identify specific acoustic properties of maternal infant-directed speech that predict physiological and behavioral calming in stressed infants, providing quantitative evidence for bidirectional biobehavioral co-regulation.
Limits
The study is an observational laboratory cohort design, which cannot establish causality. The sample size was modest (94 dyads) and restricted to biological mothers and 4–8-month-old infants, with no data on non-maternal caregivers or long-term regulatory outcomes.
Cited by
- supports Research shows that when an infant is distressed, exposure to a mother's melodic voice causes an immediate drop in heart rate, whereas a non-melodic voice does not.