揭示语音识别系统在欺骗攻击下的脆弱性,提出高效生成干扰音频的方法。
Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
- 采用快速梯度符号法与零阶优化,实现低成本白盒攻击
- 仅用35dB信噪比扰动,1分钟内生成有效对抗样本
- 验证污染攻击可使先进模型误识别音频信号
近期研究揭示了自动语音识别系统对对抗样本的脆弱性,这些样本可误导系统错误解析语音指令。尽管此前工作主要聚焦于受限优化的白盒攻击及针对商用设备的迁移性黑盒攻击,本文探索了成本低廉的白盒攻击与非迁移性黑盒对抗攻击,借鉴了快速梯度符号法和零阶优化等方法。论文新贡献在于揭示了污染攻击如何损害先进模型性能,导致音频信号误判。实验表明,混合模型可在极小扰动下生成显著影响的对抗样本,信噪比达35dB,且生成时间不足一分钟。这些开源模型的安全漏洞具有实际威胁,凸显对抗防御的重要性。
原文摘要 · Abstract (English)
Recent studies have demonstrated the vulnerability of Automatic Speech Recognition systems to adversarial examples, which can deceive these systems into misinterpreting input speech commands. While previous research has primarily focused on white-box attacks with constrained optimizations, and transferability based black-box attacks against commercial Automatic Speech Recognition devices, this paper explores cost efficient white-box attack and non transferability black-box adversarial attacks on Automatic Speech Recognition systems, drawing insights from approaches such as Fast Gradient Sign Method and Zeroth-Order Optimization. Further, the novelty of the paper includes how poisoning attack can degrade the performances of state-of-the-art models leading to misinterpretation of audio signals. Through experimentation and analysis, we illustrate how hybrid models can generate subtle yet impactful adversarial examples with very little perturbation having Signal Noise Ratio of 35dB that can be generated within a minute. These vulnerabilities of state-of-the-art open source model have practical security implications, and emphasize the need for adversarial security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。