arXiv:2504.16474cs.CRcs.LG2025-04TPAMI被引 1

通过寻找多样代理模型上的平坦极小值,提升对抗样本迁移性。

Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation

论文配图:Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
图 1 · 摘自论文原文
  • 在多种代理模型上优化对抗样本的平坦极小值,增强迁移能力。
  • 在NIPS2017和CIFAR-10上对多个目标模型攻击成功率超90%。
  • 适合研究对抗攻击迁移机制或设计鲁棒防御的读者。

基于迁移的黑盒对抗攻击需在已知代理模型上生成对抗样本(AE),使其对未知目标模型仍有效。现有方法多为启发式设计,缺乏理论支撑。本文推导出新的可迁移性界,提供可证明的保证。理论分析表明,在控制代理-目标模型差异(以对抗模型差异衡量)的前提下,使对抗样本在代理模型集上趋向平坦极小值,能全面保障迁移性。该结果催生了一般化攻击框架,揭示以往方法仅考虑部分影响因素。算法上,我们构建具有多样化对抗脆弱性的代理模型集,以缩小对抗模型差异,并提出模型多样性兼容的反向对抗扰动(DRAP),有效提升对抗样本在多样代理模型上的平坦性。在NIPS2017和CIFAR-10数据集上,针对多种目标模型的实验验证了该方法的有效性,攻击成功率超过90%。

原文摘要 · Abstract (English)

The transfer-based black-box adversarial attack setting poses the challenge of crafting an adversarial example (AE) on known surrogate models that remain effective against unseen target models. Due to the practical importance of this task, numerous methods have been proposed to address this challenge. However, most previous methods are heuristically designed and intuitively justified, lacking a theoretical foundation. To bridge this gap, we derive a novel transferability bound that offers provable guarantees for adversarial transferability. Our theoretical analysis has the advantages of \textit{(i)} deepening our understanding of previous methods by building a general attack framework and \textit{(ii)} providing guidance for designing an effective attack algorithm. Our theoretical results demonstrate that optimizing AEs toward flat minima over the surrogate model set, while controlling the surrogate-target model shift measured by the adversarial model discrepancy, yields a comprehensive guarantee for AE transferability. The results further lead to a general transfer-based attack framework, within which we observe that previous methods consider only partial factors contributing to the transferability. Algorithmically, inspired by our theoretical results, we first elaborately construct the surrogate model set in which models exhibit diverse adversarial vulnerabilities with respect to AEs to narrow an instantiated adversarial model discrepancy. Then, a \textit{model-Diversity-compatible Reverse Adversarial Perturbation} (DRAP) is generated to effectively promote the flatness of AEs over diverse surrogate models to improve transferability. Extensive experiments on NIPS2017 and CIFAR-10 datasets against various target models demonstrate the effectiveness of our proposed attack.

对抗攻击迁移性平坦极小值代理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。