arXiv:2509.21087eess.AScs.LG2025-09被引 1

先进语音增强模型易受精心设计的对抗噪声攻击,可篡改语义。

Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?

  • 用心理声学掩蔽技术隐藏对抗噪声,使其难以察觉。
  • 攻击可使增强后语音传递完全不同的语义内容。
  • 扩散模型因随机采样机制天然具备抗攻击能力。

机器学习驱动的语音增强系统表达能力日益强大,可对输入信号进行深度修改。本文揭示了这种表达力带来的新风险:先进的语音增强模型可能遭受对抗攻击。具体而言,我们证明了经过精心设计、并被原始语音信号心理声学掩蔽的对抗噪声,可被注入系统,导致增强后的语音输出传达出完全不同的语义。实验验证了当前主流预测型语音增强模型确实可被此类攻击操控。此外,我们指出采用随机采样器的扩散模型因设计原理天然具备对这类攻击的鲁棒性。

原文摘要 · Abstract (English)

Machine learning approaches for speech enhancement are becoming increasingly expressive, enabling ever more powerful modifications of input signals. In this paper, we demonstrate that this expressiveness introduces a vulnerability: advanced speech enhancement models can be susceptible to adversarial attacks. Specifically, we show that adversarial noise, carefully crafted and psychoacoustically masked by the original input, can be injected such that the enhanced speech output conveys an entirely different semantic meaning. We experimentally verify that contemporary predictive speech enhancement models can indeed be manipulated in this way. Furthermore, we highlight that diffusion models with stochastic samplers exhibit inherent robustness to such adversarial attacks by design.

语音增强对抗攻击扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。