arXiv:2601.12304cs.CVcs.AI2026-01中稿 · ICASSP 2026被引 1

提出双阶段多样攻击框架,提升视觉语言模型对抗样本生成效果

A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models

  • 分两阶段生成文本与图像扰动,增强全局多样性
  • 黑盒攻击成功率最高提升11.17%,优于现有方法
  • 模块化设计,可兼容其他攻击方法提升迁移性

视觉-语言预训练(VLP)模型在黑盒场景下易受对抗样本攻击。现有多模态攻击存在扰动多样性不足和多阶段流程不稳定的缺陷。为此,我们提出2S-GDA,一种双阶段全局多样性攻击框架。首先通过候选文本扩展与全局感知替换实现文本扰动的全局多样性;其次利用多尺度缩放与块混洗旋转生成图像级扰动以增强视觉多样性。在VLP模型上的大量实验表明,2S-GDA在黑盒设置下攻击成功率相比先进方法最高提升11.17%。该框架具有模块化特性,可轻松与其他方法结合,进一步提升对抗迁移能力。

原文摘要 · Abstract (English)

Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage pipelines. To address these challenges, we propose 2S-GDA, a two-stage globally-diverse attack framework. The proposed method first introduces textual perturbations through a globally-diverse strategy by combining candidate text expansion with globally-aware replacement. To enhance visual diversity, image-level perturbations are generated using multi-scale resizing and block-shuffle rotation. Extensive experiments on VLP models demonstrate that 2S-GDA consistently improves attack success rates over state-of-the-art methods, with gains of up to 11.17\% in black-box settings. Our framework is modular and can be easily combined with existing methods to further enhance adversarial transferability.

对抗攻击视觉语言黑盒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。