四款主流语音降噪模型可被隐蔽噪声干扰,导致输出完全听不懂。
Are Deep Speech Denoising Models Robust to Adversarial Noise?
- 通过心理声学隐藏的对抗噪声,破坏语音降噪模型输出。
- 攻击后音频在专家测试中完全无法理解,但噪声基本听不到。
- 提醒安全关键场景需警惕降噪模型的脆弱性,适合安全与语音研究者关注。
深度语音降噪(DNS)模型广泛应用于诸多高风险语音场景。然而我们发现,四种近期的DNS模型均可通过添加心理声学隐藏的对抗噪声,被诱导输出无意义的杂音,即便在低背景噪声和模拟空中传输环境下也如此。对其中三种模型的音频与多媒体专家小规模转录测试表明,受攻击音频几乎无法理解;同时,ABX测试显示对抗噪声普遍难以察觉,尽管个体与样本间存在差异。尽管我们也验证了针对攻击和模型迁移的若干负面结果,但这些发现仍凸显出在开源DNS系统用于安全关键应用前,亟需实际防护措施。
原文摘要 · Abstract (English)
Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, we show that four recent DNS models can each be reduced to outputting unintelligible gibberish through the addition of psychoacoustically hidden adversarial noise, even in low-background-noise and simulated over-the-air settings. For three of the models, a small transcription study with audio and multimedia experts confirms unintelligibility of the attacked audio; simultaneously, an ABX study shows that the adversarial noise is generally imperceptible, with some variance between participants and samples. While we also establish several negative results around targeted attacks and model transfer, our results nevertheless highlight the need for practical countermeasures before open-source DNS systems can be used in safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。