攻击语音识别系统新方法:从波形转为特征空间,更难防御且泛化更强。
Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

- 在自监督特征空间扰动,避开波形级防御
- 仅用公开模型训练,黑盒攻击使错误率提升26.6%
- 对多种防御策略仍有效,错误率提升36.2%
自动语音识别(ASR)系统广泛用于多语言语音转写,其对抗鲁棒性成为研究热点。现有攻击直接在音频波形上添加噪声,但存在两大局限:迁移性差,且易被针对输入空间扰动的防御所缓解。本文提出一种基于代理的黑盒攻击——干净参考特征-声码器攻击(Clean-Referenced Feature-Vocoder Attack),将对抗搜索空间从原始波形转移到自监督学习(SSL)表示。通过扰动更具泛化性的音素声学特征,减少对特定代理模型波形梯度的依赖,提升跨模型迁移能力。同时,将对抗信号从显式的波形噪声转移至SSL特征空间扰动,并通过声码器重构为类语音波形攻击样本,使其与依赖波形边界的防御不匹配。大量实验表明,仅以公开的Whisper-small作为代理模型优化,该攻击在黑盒ASR系统上实现+26.6 WER提升,优于当前最优基线;在多种训练防御下仍保持效果,错误率提升达+36.2。结果揭示了当前ASR鲁棒性评估中的盲点。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription. Their robustness to adversarial attacks has become an important topic for the community. Existing adversarial attacks directly add adversarial noise to the speech audio. However, prior work has shown that existing adversarial attacks face two limitations: they often transfer poorly to black-box ASR systems and are increasingly mitigated by defenses tailored to input-space perturbations. In this work, we propose a Clean-Referenced Feature-Vocoder Attack, a surrogate-based black-box attack that moves the adversarial search space from raw waveforms to self-supervised learning (SSL) representations. To address the transferability limitation, we perturb more generalizable acoustic-phonetic representations rather than low-level waveform samples, reducing dependence on surrogate-specific waveform gradients and encouraging adversarial perturbations that generalize across ASR systems. To bypass different defenses, we shift the adversarial signal from explicit additive waveform noise to SSL feature-space perturbations and reconstruct them through a vocoder into speech-like waveform adversarial signals, making the resulting samples less aligned with waveform-bounded defenses. Extensive experiments show that, when optimized only on raw Whisper-small as a public surrogate model, our attack transfers effectively to black-box ASR models with a +26.6 WER improvement over the SOTA baseline, while also remaining effective against multiple training defenses with a +36.2 WER improvement. These results reveal a blind spot in current ASR robustness evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。