攻击者仅修改1.5%语音数据,即可让智能助手误听指令。
Clean Label Attacks against SLU Systems
- 不改标签只藏触发词,悄悄污染训练数据
- 仅污染1.5%数据就实现99.3%攻击成功率
- 针对难辨认样本攻击更有效,适合研究安全漏洞
中毒后门攻击通过操纵训练数据,在推理时插入触发信号诱导目标模型产生特定行为。本文将无标签污染的干净标签后门(CLBD)攻击应用于先进的语音识别模型(支持/执行语音语言理解任务),在仅污染10%训练数据的情况下,达到99.8%的攻击成功率。分析了毒化信号强度、污染比例及触发器选择对攻击的影响,发现当攻击集中在代理模型难以识别的样本上时效果最佳,仅污染1.5%数据即可实现99.3%的成功率。此外,测试了两种针对梯度攻击的防御方法,发现其对中毒攻击效果有限。
原文摘要 · Abstract (English)
Poisoning backdoor attacks involve an adversary manipulating the training data to induce certain behaviors in the victim model by inserting a trigger in the signal at inference time. We adapted clean label backdoor (CLBD)-data poisoning attacks, which do not modify the training labels, on state-of-the-art speech recognition models that support/perform a Spoken Language Understanding task, achieving 99.8% attack success rate by poisoning 10% of the training data. We analyzed how varying the signal-strength of the poison, percent of samples poisoned, and choice of trigger impact the attack. We also found that CLBD attacks are most successful when applied to training samples that are inherently hard for a proxy model. Using this strategy, we achieved an attack success rate of 99.3% by poisoning a meager 1.5% of the training data. Finally, we applied two previously developed defenses against gradient-based attacks, and found that they attain mixed success against poisoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。