Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding
Level 5 - mechanism / opinion, no new human data
Level 5 by design analogy (computational algorithm design and bench evaluation; non-clinical CEBM)
OpenAlex W2885407017 · doi:10.14722/ndss.2019.23288
What was done
The authors developed a targeted adversarial attack against deep neural network (DNN)-based automatic speech recognition (ASR) systems using psychoacoustic hiding. The method uses an additional backpropagation step to learn adversarial perturbations beneath human auditory perception thresholds via a psychoacoustic model, combined with forced alignment to optimize temporal alignment between carrier audio and malicious target text. The attack was experimentally evaluated against the Kaldi ASR system, and human user studies were conducted to assess the perceptibility of the perturbations.
What was found
The attack succeeded in up to 98% of evaluated cases against Kaldi, requiring less than two minutes of computational time for a 10-second audio sample. In user perceptual evaluations, none of the target malicious transcriptions were audible to human listeners, and listeners comprehended the original speech content with unchanged accuracy. Specific numerical data regarding participant numbers, trial counts, and exact comprehension scores were not reported in the abstract.
Why it matters
This work demonstrates that voice-controlled systems can be manipulated with targeted malicious commands hidden entirely beneath human auditory perception thresholds. It highlights a critical security vulnerability in modern ASR pipelines arising from differences between human psychoacoustics and machine feature extraction.
Limits
The abstract does not disclose the sample size, the number of user study participants, or specific audio datasets tested. The evaluation was restricted to a single ASR framework (Kaldi) in a computational environment, with no testing reported for over-the-air acoustic playback, ambient noise conditions, or physical microphone distortions.
Cited by
- supports Perceptual lossy audio compression formats like MP3 utilize psychoacoustic models of auditory masking, discarding sounds that the human auditory system cannot perceive when played simultaneously with louder sounds at adjacent frequencies.