提出精准攻击方法M2A,让声音事件检测系统只误判目标区域。
Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems
- 通过保留非目标区域输出,实现精准定位攻击
- 在两个先进模型上达到94.56%和99.11%的编辑精度
- 适合研究音频安全与对抗攻防的学者
声音事件检测(SED)系统正被广泛应用于工业监控和音频安防等关键场景。然而,其对对抗攻击的鲁棒性尚未得到充分研究。现有针对具备检测与定位能力的多音事件检测系统的音频对抗攻击,常因SED强上下文依赖而效果不佳,或仅关注将目标区域错误分类为特定事件,导致非目标区域也被无意干扰。为此,我们提出镜像与静音攻击(M2A)框架,专为多音事件检测系统设计定向对抗攻击。在优化过程中,我们对非目标输出施加特定约束,称为保持损失(preservation loss),确保攻击不改变非目标区域的模型输出,从而实现精准攻击。此外,我们引入一种新评估指标——编辑精度(Editing Precision, EP),平衡攻击效果与精度,使方法同时提升两者。大量实验表明,M2A在两个先进SED模型上分别达到94.56%和99.11%的EP,证明该框架既有效又显著提升了攻击精度。
原文摘要 · Abstract (English)
Sound Event Detection (SED) systems are increasingly deployed in safety-critical applications such as industrial monitoring and audio surveillance. However, their robustness against adversarial attacks has not been well explored. Existing audio adversarial attacks targeting SED systems, which incorporate both detection and localization capabilities, often lack effectiveness due to SED's strong contextual dependencies or lack precision by focusing solely on misclassifying the target region as the target event, inadvertently affecting non-target regions. To address these challenges, we propose the Mirage and Mute Attack (M2A) framework, which is designed for targeted adversarial attacks on polyphonic SED systems. In our optimization process, we impose specific constraints on the non-target output, which we refer to as preservation loss, ensuring that our attack does not alter the model outputs for non-target region, thus achieving precise attacks. Furthermore, we introduce a novel evaluation metric Editing Precison (EP) that balances effectiveness and precision, enabling our method to simultaneously enhance both. Comprehensive experiments show that M2A achieves 94.56% and 99.11% EP on two state-of-the-art SED models, demonstrating that the framework is sufficiently effective while significantly enhancing attack precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。