解析模型集成攻击的可迁移性原理,给出提升攻击效果的实用方法。
Understanding Model Ensemble in Transferable Adversarial Attack
- 定义迁移误差,分解为脆弱性与多样性因素
- 实验证明增加代理模型数量和多样性可降低误差
- 适合研究对抗样本迁移性或模型安全的学者参考
模型集成对抗攻击已成为生成可迁移对抗样本的强大方法,能针对未知模型发起攻击,但其理论基础仍不清晰。本文首次提出迁移误差概念,结合多样性和经验模型集成雷达马克复杂度,将迁移误差分解为脆弱性、多样性与常数项,严格解释了模型集成攻击中迁移误差的来源:对抗样本对集成组件的脆弱性及其组件间的多样性。进一步运用信息论最新数学工具,通过复杂度与泛化项界定了迁移误差,提出三条降低误差的实用指导:(1)引入更多代理模型;(2)增强其多样性;(3)在过拟合时降低模型复杂度。在54个模型上的大量实验验证了理论框架的有效性,标志着对可迁移模型集成对抗攻击理解的重要进展。
原文摘要 · Abstract (English)
Model ensemble adversarial attack has become a powerful method for generating transferable adversarial examples that can target even unknown models, but its theoretical foundation remains underexplored. To address this gap, we provide early theoretical insights that serve as a roadmap for advancing model ensemble adversarial attack. We first define transferability error to measure the error in adversarial transferability, alongside concepts of diversity and empirical model ensemble Rademacher complexity. We then decompose the transferability error into vulnerability, diversity, and a constant, which rigidly explains the origin of transferability error in model ensemble attack: the vulnerability of an adversarial example to ensemble components, and the diversity of ensemble components. Furthermore, we apply the latest mathematical tools in information theory to bound the transferability error using complexity and generalization terms, contributing to three practical guidelines for reducing transferability error: (1) incorporating more surrogate models, (2) increasing their diversity, and (3) reducing their complexity in cases of overfitting. Finally, extensive experiments with 54 models validate our theoretical framework, representing a significant step forward in understanding transferable model ensemble adversarial attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。