A high-performance neuroprosthesis for speech decoding and avatar control.
Level 4 - case-series / case-control
Single-participant interventional case study / proof-of-concept trial
PubMed 37612505 · doi:10.1038/s41586-023-06443-4
What was done
High-density surface recordings of the speech cortex were collected from a single clinical trial participant with severe limb and vocal paralysis during attempted silent speech. Deep-learning models were trained to decode neural activity in real time across three output modalities: text, synthesized speech audio personalized to the participant's pre-injury voice, and facial-avatar animation for speech and non-speech gestures.
What was found
Text decoding achieved a median rate of 78 words per minute with a median word error rate of 25% on a large vocabulary. The decoders achieved functional performance across text, voice synthesis, and avatar animation with less than two weeks of model training.
Why it matters
This study demonstrates that multimodal speech neuroprostheses can decode silent speech into text, synthesized voice, and expressive avatar animations at speeds substantially faster than earlier communication interfaces.
Limits
Evaluated in only one participant (n = 1), so generalizability across different etiologies of paralysis, brain injury locations, and user baselines is unknown. Long-term decoder stability and performance outside experimental settings were not reported.
Cited by
- supports Neuroengineering research led by Dr. Edward Chang has mapped neural activity to vocal tract control (larynx and pharynx) to decode speech and enable communication in paralyzed patients.