arXiv:2410.04884cs.CVcs.AI2024-10中稿 · Visual Intelligenc…被引 56

用自然图像块攻击视觉语言模型,不改文字也能100%骗过。

Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models

  • 仅用图像块扰动,不修改文本内容,更隐蔽。
  • 结合扩散模型生成逼真噪声,攻击成功率100%。
  • 适合研究多模态安全或对抗样本的学者。

视觉语言预训练(VLP)模型在多个领域表现优异,但易受对抗攻击。传统方法需同时扰动图像和文本,存在现实迁移性差、文本修改明显等问题。本文提出仅使用图像块进行攻击的新策略,保持原文本不变。利用扩散模型先验增强扰动的真实性与自然性,并通过跨注意力机制生成注意力图,指导最优补丁位置。白盒环境下,图像到文本任务中攻击成功率达100%;在文本到图像迁移任务中也表现出色。

原文摘要 · Abstract (English)

Visual language pre-training (VLP) models have demonstrated significant success across various domains, yet they remain vulnerable to adversarial attacks. Addressing these adversarial vulnerabilities is crucial for enhancing security in multimodal learning. Traditionally, adversarial methods targeting VLP models involve simultaneously perturbing images and text. However, this approach faces notable challenges: first, adversarial perturbations often fail to translate effectively into real-world scenarios; second, direct modifications to the text are conspicuously visible. To overcome these limitations, we propose a novel strategy that exclusively employs image patches for attacks, thus preserving the integrity of the original text. Our method leverages prior knowledge from diffusion models to enhance the authenticity and naturalness of the perturbations. Moreover, to optimize patch placement and improve the efficacy of our attacks, we utilize the cross-attention mechanism, which encapsulates intermodal interactions by generating attention maps to guide strategic patch placements. Comprehensive experiments conducted in a white-box setting for image-to-text scenarios reveal that our proposed method significantly outperforms existing techniques, achieving a 100% attack success rate. Additionally, it demonstrates commendable performance in transfer tasks involving text-to-image configurations.

对抗攻击多模态视觉语言图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。