提出PBCAT方法,用局部和全局对抗扰动提升目标检测器对物理攻击的防御力。
PBCAT: Patch-based composite adversarial training against physically realizable attacks on object detection
- 融合小区域梯度引导的对抗贴纸与全图不可察觉的扰动进行训练
- 在新型对抗纹理攻击下,检测准确率比之前方法提升29.7%
- 可防御多种未见过的物理可实现攻击,适合安全敏感场景
目标检测在众多安全敏感应用中至关重要。然而,近期研究表明,目标检测器易受物理可实现攻击(如对抗贴纸和对抗纹理)欺骗,构成现实且紧迫的威胁。对抗训练(AT)被公认为最有效的防御手段。尽管在分类模型的 $l_ty$ 攻击设置下已有广泛研究,针对目标检测器的物理可实现攻击的对抗训练仍探索不足。早期工作仅聚焦于对抗贴纸防御,对更广泛的物理攻击缺乏系统研究。本文提出一种统一的对抗训练方法——PBCAT(Patch-Based Composite Adversarial Training)。PBCAT通过结合小区域梯度引导的对抗贴纸与覆盖整图的不可察觉全局扰动来优化模型。该设计使其不仅能防御对抗贴纸,还可抵御未见的物理可实现攻击,如对抗纹理。大量实验表明,相较于现有先进防御方法,PBCAT在多种设置下显著提升了鲁棒性。特别地,在一种新型对抗纹理攻击下,检测准确率提升了29.7%。
原文摘要 · Abstract (English)
Object detection plays a crucial role in many security-sensitive applications. However, several recent studies have shown that object detectors can be easily fooled by physically realizable attacks, \eg, adversarial patches and recent adversarial textures, which pose realistic and urgent threats. Adversarial Training (AT) has been recognized as the most effective defense against adversarial attacks. While AT has been extensively studied in the $l_\infty$ attack settings on classification models, AT against physically realizable attacks on object detectors has received limited exploration. Early attempts are only performed to defend against adversarial patches, leaving AT against a wider range of physically realizable attacks under-explored. In this work, we consider defending against various physically realizable attacks with a unified AT method. We propose PBCAT, a novel Patch-Based Composite Adversarial Training strategy. PBCAT optimizes the model by incorporating the combination of small-area gradient-guided adversarial patches and imperceptible global adversarial perturbations covering the entire image. With these designs, PBCAT has the potential to defend against not only adversarial patches but also unseen physically realizable attacks such as adversarial textures. Extensive experiments in multiple settings demonstrated that PBCAT significantly improved robustness against various physically realizable attacks over state-of-the-art defense methods. Notably, it improved the detection accuracy by 29.7\% over previous defense methods under one recent adversarial texture attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。