arXiv:2511.00831cs.CVcs.AI2025-11NAACL被引 1

通过局部打乱与采样增强多模态对抗样本的迁移能力

Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack

  • 对图像局部块随机打乱并采样,扩充输入多样性
  • 在多个模型和数据集上显著提升对抗样本迁移率
  • 适合研究对抗鲁棒性或攻击生成的学者使用

视觉-语言预训练(VLP)模型在多种下游任务中表现优异,但仍易受对抗样本攻击。现有方法虽通过跨模态交互提升多模态对抗样本的迁移性,但因过度依赖单一模态的对抗信息,导致过拟合问题。为此,本文受部分对抗训练策略启发,提出一种新攻击方法——局部打乱与样本采样攻击(LSSA)。LSSA随机打乱图像的一个局部块,扩展原始图像-文本对,生成对抗图像,并在附近采样;再结合原始与采样图像生成对抗文本。大量实验表明,LSSA显著提升多模态对抗样本在不同VLP模型及下游任务中的迁移能力,且优于其他先进攻击方法,在大型视觉-语言模型上表现更优。

原文摘要 · Abstract (English)

Visual-Language Pre-training (VLP) models have achieved significant performance across various downstream tasks. However, they remain vulnerable to adversarial examples. While prior efforts focus on improving the adversarial transferability of multimodal adversarial examples through cross-modal interactions, these approaches suffer from overfitting issues, due to a lack of input diversity by relying excessively on information from adversarial examples in one modality when crafting attacks in another. To address this issue, we draw inspiration from strategies in some adversarial training methods and propose a novel attack called Local Shuffle and Sample-based Attack (LSSA). LSSA randomly shuffles one of the local image blocks, thus expanding the original image-text pairs, generating adversarial images, and sampling around them. Then, it utilizes both the original and sampled images to generate the adversarial texts. Extensive experiments on multiple models and datasets demonstrate that LSSA significantly enhances the transferability of multimodal adversarial examples across diverse VLP models and downstream tasks. Moreover, LSSA outperforms other advanced attacks on Large Vision-Language Models.

对抗攻击多模态视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。