arXiv:2411.02669cs.CV2024-11TPAMI被引 26

通过语义对齐三角进化生成高迁移性视觉语言攻击

Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack

论文配图:Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack
图 1 · 摘自论文原文
  • 构建干净样本、历史与当前对抗样本的三角演化空间,增强多样性
  • 在语义特征对比空间中生成攻击,提升跨模型迁移成功率
  • 适合研究视觉语言模型安全与鲁棒性的研究人员

视觉语言预训练(VLP)模型在图像与文本理解方面表现优异,但易受多模态对抗样本攻击。提升对抗样本在未见模型间的迁移能力是增强VLP模型鲁棒性与实用性的关键。现有方法通过扩充图像-文本对来增加生成过程中的多样性,以扩大图像-文本特征的对比空间,从而提升迁移性,但仅关注当前对抗样本周围的多样性,提升有限。为此,我们提出利用对抗优化轨迹中的交集区域来提升对抗样本多样性。具体地,从由原始样本、历史样本和当前对抗样本构成的对抗演化三角中采样,以增强对抗多样性。我们提供了理论分析证明该方法的有效性。此外,发现冗余非活跃维度会主导相似性计算,扭曲特征匹配,导致对抗样本依赖特定模型、迁移性下降。因此,我们提出在语义图像-文本特征对比空间中生成对抗样本,可将原始特征空间投影至语义语料子空间。该语义对齐子空间能减少图像特征冗余,从而提升对抗迁移性。在多个数据集与模型上的大量实验表明,所提方法能有效提升对抗迁移性,优于现有最先进攻击方法。代码已开源。

原文摘要 · Abstract (English)

Vision-language pre-training (VLP) models excel at interpreting both images and text but remain vulnerable to multimodal adversarial examples (AEs). Advancing the generation of transferable AEs, which succeed across unseen models, is key to developing more robust and practical VLP models. Previous approaches augment image-text pairs to enhance diversity within the adversarial example generation process, aiming to improve transferability by expanding the contrast space of image-text features. However, these methods focus solely on diversity around the current AEs, yielding limited gains in transferability. To address this issue, we propose to increase the diversity of AEs by leveraging the intersection regions along the adversarial trajectory during optimization. Specifically, we propose sampling from adversarial evolution triangles composed of clean, historical, and current adversarial examples to enhance adversarial diversity. We provide a theoretical analysis to demonstrate the effectiveness of the proposed adversarial evolution triangle. Moreover, we find that redundant inactive dimensions can dominate similarity calculations, distorting feature matching and making AEs model-dependent with reduced transferability. Hence, we propose to generate AEs in the semantic image-text feature contrast space, which can project the original feature space into a semantic corpus subspace. The proposed semantic-aligned subspace can reduce the image feature redundancy, thereby improving adversarial transferability. Extensive experiments across different datasets and models demonstrate that the proposed method can effectively improve adversarial transferability and outperform state-of-the-art adversarial attack methods. The code is released at https://github.com/jiaxiaojunQAQ/SA-AET.

对抗攻击视觉语言迁移性特征对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。