提出新方法提升物理世界对抗样本的攻击效果与隐蔽性。
Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust Feature
- 利用鲁棒特征覆盖提升对抗样本的迁移性和环境鲁棒性。
- 通过去除冗余扰动,保留关键对抗模式,增强隐蔽性。
- 适用于复杂模型如视觉语言大模型,适用范围广。
随着深度神经网络在现实世界中的广泛应用,物理世界对抗样本(PAEs)成为研究热点,其通过引入扰动导致模型输出错误。然而现有方法面临两大挑战:攻击性能不佳(迁移性差、对环境变化不鲁棒),以及难以平衡攻击效果与隐蔽性——攻击越强,越容易被察觉。本文提出一种基于扰动的新方法,针对第一挑战,设计欺骗性鲁棒特征注入策略,利用对扰动具有预测性且跨模型一致的鲁棒特征(RFs),将其他类别的鲁棒特征覆盖到干净图像的预测特征上,显著提升迁移性和鲁棒性。针对第二挑战,提出对抗语义模式最小化策略,移除大部分扰动,仅保留必要对抗模式。据此构建鲁棒特征覆盖攻击(RFCoA),包含鲁棒特征解耦与对抗特征融合两阶段:先在特征空间提取目标类别鲁棒特征,再通过注意力机制融合这些特征至干净图像的预测特征,并清除冗余扰动。实验表明,该方法在迁移性、鲁棒性和隐蔽性上均优于现有最先进方法。此外,其有效性可扩展至大型视觉-语言模型(LVLMs),表明其在更复杂任务中的潜力。
原文摘要 · Abstract (English)
As deep neural networks (DNNs) are widely applied in the physical world, many researches are focusing on physical-world adversarial examples (PAEs), which introduce perturbations to inputs and cause the model's incorrect outputs. However, existing PAEs face two challenges: unsatisfactory attack performance (i.e., poor transferability and insufficient robustness to environment conditions), and difficulty in balancing attack effectiveness with stealthiness, where better attack effectiveness often makes PAEs more perceptible. In this paper, we explore a novel perturbation-based method to overcome the challenges. For the first challenge, we introduce a strategy Deceptive RF injection based on robust features (RFs) that are predictive, robust to perturbations, and consistent across different models. Specifically, it improves the transferability and robustness of PAEs by covering RFs of other classes onto the predictive features in clean images. For the second challenge, we introduce another strategy Adversarial Semantic Pattern Minimization, which removes most perturbations and retains only essential adversarial patterns in AEsBased on the two strategies, we design our method Robust Feature Coverage Attack (RFCoA), comprising Robust Feature Disentanglement and Adversarial Feature Fusion. In the first stage, we extract target class RFs in feature space. In the second stage, we use attention-based feature fusion to overlay these RFs onto predictive features of clean images and remove unnecessary perturbations. Experiments show our method's superior transferability, robustness, and stealthiness compared to existing state-of-the-art methods. Additionally, our method's effectiveness can extend to Large Vision-Language Models (LVLMs), indicating its potential applicability to more complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。