Lea Schönherr · 2019 · computational experiment and perceptual user study · n=?

Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic Hiding

Cited 49 times in the scientific literature.

Level 5 - mechanism / opinion, no new human data

Level 5 by design analogy (computational algorithm design and bench evaluation; non-clinical CEBM)

OpenAlex W2885407017 · doi:10.14722/ndss.2019.23288 · record verified 2026-08-26

What was done

The authors developed a targeted adversarial attack against deep neural network (DNN)-based automatic speech recognition (ASR) systems using psychoacoustic hiding. The method uses an additional backpropagation step to learn adversarial perturbations beneath human auditory perception thresholds via a psychoacoustic model, combined with forced alignment to optimize temporal alignment between carrier audio and malicious target text. The attack was experimentally evaluated against the Kaldi ASR system, and human user studies were conducted to assess the perceptibility of the perturbations.

What was found

The attack succeeded in up to 98% of evaluated cases against Kaldi, requiring less than two minutes of computational time for a 10-second audio sample. In user perceptual evaluations, none of the target malicious transcriptions were audible to human listeners, and listeners comprehended the original speech content with unchanged accuracy. Specific numerical data regarding participant numbers, trial counts, and exact comprehension scores were not reported in the abstract.

Why it matters

This work demonstrates that voice-controlled systems can be manipulated with targeted malicious commands hidden entirely beneath human auditory perception thresholds. It highlights a critical security vulnerability in modern ASR pipelines arising from differences between human psychoacoustics and machine feature extraction.

Limits

The abstract does not disclose the sample size, the number of user study participants, or specific audio datasets tested. The evaluation was restricted to a single ASR framework (Kaldi) in a computational environment, with no testing reported for over-the-air acoustic playback, ambient noise conditions, or physical microphone distortions.

Cited by