噪声防御可能被强化学习攻击利用,反而提升攻击成功率。
Noise as a Double-Edged Sword: Reinforcement Learning Exploits Randomized Defenses in Neural Networks
- 用强化学习攻击噪声防御机制,发现其可被反向利用。
- 在部分视觉噪声类别上,攻击成功率提升最高达20%。
- 适合关注对抗攻击与防御安全性的研究人员参考。
本研究探讨了对抗机器学习中一个反直觉现象:基于噪声的防御策略在某些情况下可能意外助长规避攻击。尽管随机性常被用于抵御对抗样本,我们的研究发现,当面对使用强化学习(RL)的自适应攻击者时,该方法有时会适得其反。结果显示,在特定场景下,尤其是视觉噪声较大的类别中,分类器置信度引入噪声后,可被RL攻击者利用,导致规避成功率显著上升。在某些情况下,噪声防御策略相比其他方法提升了高达20%的攻击成功率。然而,这一效应并非在所有测试分类器中一致出现,凸显了噪声防御与不同模型之间复杂交互关系。结果表明,噪声防御可能无意中形成对RL攻击者有利的对抗训练循环。研究强调需采用更精细的防御策略,尤其在安全关键应用中。它挑战了‘随机性普遍增强防御’的假设,突显设计鲁棒防御机制时必须考虑自适应的强化学习攻击者。
原文摘要 · Abstract (English)
This study investigates a counterintuitive phenomenon in adversarial machine learning: the potential for noise-based defenses to inadvertently aid evasion attacks in certain scenarios. While randomness is often employed as a defensive strategy against adversarial examples, our research reveals that this approach can sometimes backfire, particularly when facing adaptive attackers using reinforcement learning (RL). Our findings show that in specific cases, especially with visually noisy classes, the introduction of noise in the classifier's confidence values can be exploited by the RL attacker, leading to a significant increase in evasion success rates. In some instances, the noise-based defense scenario outperformed other strategies by up to 20\% on a subset of classes. However, this effect was not consistent across all classifiers tested, highlighting the complexity of the interaction between noise-based defenses and different models. These results suggest that in some cases, noise-based defenses can inadvertently create an adversarial training loop beneficial to the RL attacker. Our study emphasizes the need for a more nuanced approach to defensive strategies in adversarial machine learning, particularly in safety-critical applications. It challenges the assumption that randomness universally enhances defense against evasion attacks and highlights the importance of considering adaptive, RL-based attackers when designing robust defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。