针对双人同时说话场景,提出可选择性屏蔽的语音对抗攻击方法。
Selective Masking Adversarial Attack on Automatic Speech Recognition Systems
- 通过选择性掩蔽实现双语音源中目标语音识别,另一语音静音。
- 在Conformer-CTC上实现100%攻击成功率,信噪比达37.15dB。
- 适用于研究语音系统安全性的研究人员与防御设计者。
大量研究表明自动语音识别(ASR)系统易受音频对抗攻击。现有攻击多聚焦单源场景,忽视两人同时说话的双源场景。为此,我们提出选择性掩蔽对抗攻击(SMA攻击),在双源场景下确保一个音频源被识别,而另一音频源被静音。为更好适配双源场景,SMA攻击从静音音频和选定音频构建正常双源音频。攻击初始化时使用小量高斯噪声生成对抗扰动,并通过选择性掩蔽优化算法迭代优化。大量实验表明,SMA攻击可在双源场景生成有效且不可察觉的对抗音频样本,在Conformer-CTC上平均攻击成功率达100%,信噪比为37.15dB,优于基线方法。
原文摘要 · Abstract (English)
Extensive research has shown that Automatic Speech Recognition (ASR) systems are vulnerable to audio adversarial attacks. Current attacks mainly focus on single-source scenarios, ignoring dual-source scenarios where two people are speaking simultaneously. To bridge the gap, we propose a Selective Masking Adversarial attack, namely SMA attack, which ensures that one audio source is selected for recognition while the other audio source is muted in dual-source scenarios. To better adapt to the dual-source scenario, our SMA attack constructs the normal dual-source audio from the muted audio and selected audio. SMA attack initializes the adversarial perturbation with a small Gaussian noise and iteratively optimizes it using a selective masking optimization algorithm. Extensive experiments demonstrate that the SMA attack can generate effective and imperceptible audio adversarial examples in the dual-source scenario, achieving an average success rate of attack of 100% and signal-to-noise ratio of 37.15dB on Conformer-CTC, outperforming the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。