用知识蒸馏提升对抗补丁的隐蔽性与攻击效果
Distillation-Enhanced Physical Adversarial Attacks
- 通过知识蒸馏将强攻击补丁的知识迁移到隐蔽补丁
- 攻击成功率提升20%,同时保持良好隐蔽性
- 适合研究对抗样本防御或实际场景安全评估者
物理对抗补丁的研究对揭示基于AI识别系统的漏洞、提升深度学习模型鲁棒性至关重要。尽管近期工作聚焦于增强补丁的隐蔽性以提高实用性,但在隐蔽性与攻击性能之间实现有效平衡仍是重大挑战。为此,本文提出一种新型物理对抗攻击方法,利用知识蒸馏技术。首先,针对目标环境定义一种隐蔽色彩空间以实现自然融合;其次,在无约束色彩空间中优化一个作为‘教师’的对抗补丁;最后,通过对抗知识蒸馏模块将教师补丁的知识迁移至‘学生’补丁,指导隐蔽补丁的优化。实验表明,该方法在保持隐蔽性的前提下,攻击性能提升20%,凸显其实际应用价值。
原文摘要 · Abstract (English)
The study of physical adversarial patches is crucial for identifying vulnerabilities in AI-based recognition systems and developing more robust deep learning models. While recent research has focused on improving patch stealthiness for greater practical applicability, achieving an effective balance between stealth and attack performance remains a significant challenge. To address this issue, we propose a novel physical adversarial attack method that leverages knowledge distillation. Specifically, we first define a stealthy color space tailored to the target environment to ensure smooth blending. Then, we optimize an adversarial patch in an unconstrained color space, which serves as the 'teacher' patch. Finally, we use an adversarial knowledge distillation module to transfer the teacher patch's knowledge to the 'student' patch, guiding the optimization of the stealthy patch. Experimental results show that our approach improves attack performance by 20%, while maintaining stealth, highlighting its practical value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。