arXiv:2607.15724cs.CRcs.LG2026-07ICML被引 20

用日常自然声音做触发器,实现隐蔽高效的语音识别后门攻击

Natural Backdoor Attacks on Speech Recognition Models

论文配图:Natural Backdoor Attacks on Speech Recognition Models
图 1 · 摘自论文原文
  • 用生活中常见的声音作触发器,无需修改模型结构
  • 仅5%污染样本即可达近100%攻击成功率,且不影响正常识别性能
  • 触发器天然存在、难以察觉,适合隐蔽攻击场景

随着深度学习的快速发展,其脆弱性逐渐显现。本文聚焦于语音识别系统的后门攻击,采用日常生活中常见的自然声音作为触发器,开展自然后门攻击。在两个数据集和三个模型上验证了该攻击的有效性,并研究了污染比例、触发器持续时间和混合比例对攻击效果的影响。结果表明,自然后门攻击在不损害良性样本性能的前提下,仍能实现高攻击成功率,即使触发器持续时间短或振幅低亦然。仅需5%的污染样本即可达成接近100%的攻击成功率。此外,后门可被自然界中的对应声音自动激活,难以被检测,危害更为严重。

原文摘要 · Abstract (English)

With the rapid development of deep learning, its vulnerability has gradually emerged in recent years. This work focuses on backdoor attacks on speech recognition systems. We adopt sounds that are ordinary in nature or in our daily life as triggers for natural backdoor attacks. We conduct experiments on two datasets and three models to validate the performance of natural backdoor attacks and explore the effects of poisoning rate, trigger duration and blend ratio on the performance of natural backdoor attacks. Our results show that natural backdoor attacks have a high attack success rate without compromising model performance on benign samples, even with short or low-amplitude triggers. It requires only 5% of poisoned samples to achieve a near 100% attack success rate. In addition, the backdoor will be automatically activated by the corresponding sound in nature, which is not easy to be detected and will bring severer harm.

语音安全后门攻击自然触发器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。