提出新攻击方法,让对抗样本在更平坦区域生成,提升跨模型攻击成功率。
Beyond Deceptive Flatness: Dual-Order Solution for Strengthening Adversarial Transferability
- 利用双阶信息设计新攻击,解决虚假平坦性问题
- 在ImageNet兼容数据集上超越6个基线,提升跨架构迁移能力
- 适合研究对抗攻击迁移性的学者与安全评估人员
迁移攻击通过在代理模型上生成对抗样本,以欺骗未知的受害者模型,构成现实威胁并引发广泛关注。尽管已有研究聚焦于平坦损失以提升迁移性,但仍在次优区域,特别是所谓的‘虚假平坦’区域(flat-yet-sharp)中徘徊。本文从双阶信息角度提出一种新型黑盒梯度攻击方法。我们引入对抗平坦性(Adversarial Flatness, AF),有效应对虚假平坦问题,并提供对抗迁移性的理论保障。基于该目标的高效近似,我们实现为对抗平坦性攻击(AFA),解决了梯度符号变化问题。为进一步提升攻击性能,我们设计蒙特卡洛对抗采样(MCAS),增强内循环采样效率。在ImageNet兼容数据集上的全面实验表明,本方法优于六个基线,生成位于更平坦区域的对抗样本,并显著提升跨模型架构的迁移能力。在输入变换攻击或百度云API测试中,本方法同样表现更优。
原文摘要 · Abstract (English)
Transferable attacks generate adversarial examples on surrogate models to fool unknown victim models, posing real-world threats and growing research interest. Despite focusing on flat losses for transferable adversarial examples, recent studies still fall into suboptimal regions, especially the flat-yet-sharp areas, termed as deceptive flatness. In this paper, we introduce a novel black-box gradient-based transferable attack from a perspective of dual-order information. Specifically, we feasibly propose Adversarial Flatness (AF) to the deceptive flatness problem and a theoretical assurance for adversarial transferability. Based on this, using an efficient approximation of our objective, we instantiate our attack as Adversarial Flatness Attack (AFA), addressing the altered gradient sign issue. Additionally, to further improve the attack ability, we devise MonteCarlo Adversarial Sampling (MCAS) by enhancing the inner-loop sampling efficiency. The comprehensive results on ImageNet-compatible dataset demonstrate superiority over six baselines, generating adversarial examples in flatter regions and boosting transferability across model architectures. When tested on input transformation attacks or the Baidu Cloud API, our method outperforms baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。