提出HEAT方法,提升对抗样本跨模型迁移能力
Harmonizing Intra-coherence and Inter-divergence in Ensemble Attacks for Adversarial Transferability
- 用SVD合成模型间共享梯度方向
- 动态平衡单模型稳定性和多模型差异性
- 适用于研究对抗攻击迁移性的研究人员
模型集成攻击显著提升了对抗样本的迁移性,但也对深度神经网络安全构成严重威胁。现有方法面临两大挑战:难以捕捉模型间的共享梯度方向,缺乏自适应权重分配机制。为此,我们提出首个将领域泛化引入对抗样本生成的新方法HEAT。HEAT包含两个核心模块:共识梯度方向合成器(利用奇异值分解合成共享梯度方向),以及双和谐权重调度器(动态平衡域内一致性与域间多样性)。实验表明,HEAT在多种数据集和设置下均显著优于现有方法,为对抗攻击研究提供了新视角。
原文摘要 · Abstract (English)
The development of model ensemble attacks has significantly improved the transferability of adversarial examples, but this progress also poses severe threats to the security of deep neural networks. Existing methods, however, face two critical challenges: insufficient capture of shared gradient directions across models and a lack of adaptive weight allocation mechanisms. To address these issues, we propose a novel method Harmonized Ensemble for Adversarial Transferability (HEAT), which introduces domain generalization into adversarial example generation for the first time. HEAT consists of two key modules: Consensus Gradient Direction Synthesizer, which uses Singular Value Decomposition to synthesize shared gradient directions; and Dual-Harmony Weight Orchestrator which dynamically balances intra-domain coherence, stabilizing gradients within individual models, and inter-domain diversity, enhancing transferability across models. Experimental results demonstrate that HEAT significantly outperforms existing methods across various datasets and settings, offering a new perspective and direction for adversarial attack research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。