提出跨模态图像转视频攻击,提升对抗样本在未知视频模型间的迁移能力。
Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach
- 用图像大模型做代理,生成可跨模型迁移的视频对抗样本。
- 在多个任务上实现58%以上攻击成功率,接近白盒攻击效果。
- 适合研究视频多模态模型安全性的研究人员参考。
基于视频的多模态大语言模型(V-MLLMs)在视频-文本多模态任务中易受对抗样本攻击。然而,对抗视频在未见模型间的迁移性——这一常见且实际的黑盒场景——尚未被探索。本文首次系统研究了对抗视频在V-MLLM间的迁移性。发现现有攻击方法在黑盒设置下表现受限,主要因:(1)视频特征扰动缺乏泛化性,(2)仅关注稀疏关键帧,(3)未能融合多模态信息。为此,我们提出图像转视频多模态大模型攻击(I2V-MLLM),利用图像多模态大模型(I-MLLM)作为代理模型生成对抗视频。通过整合多模态交互与时空信息,干扰视频表示的潜在空间,提升迁移性。此外,引入扰动传播技术以应对不同未知帧采样策略。实验表明,该方法在多个视频-文本多模态任务上实现了强迁移性。相比白盒攻击,使用BLIP-2作为代理模型的黑盒攻击在零样本视频问答任务中分别达到57.98%(MSVD-QA)和58.26%(MSRVTT-QA)的平均攻击成功率达(AASR),表现相当。
原文摘要 · Abstract (English)
Video-based multimodal large language models (V-MLLMs) have shown vulnerability to adversarial examples in video-text multimodal tasks. However, the transferability of adversarial videos to unseen models - a common and practical real-world scenario - remains unexplored. In this paper, we pioneer an investigation into the transferability of adversarial video samples across V-MLLMs. We find that existing adversarial attack methods face significant limitations when applied in black-box settings for V-MLLMs, which we attribute to the following shortcomings: (1) lacking generalization in perturbing video features, (2) focusing only on sparse key-frames, and (3) failing to integrate multimodal information. To address these limitations and deepen the understanding of V-MLLM vulnerabilities in black-box scenarios, we introduce the Image-to-Video MLLM (I2V-MLLM) attack. In I2V-MLLM, we utilize an image-based multimodal large language model (I-MLLM) as a surrogate model to craft adversarial video samples. Multimodal interactions and spatiotemporal information are integrated to disrupt video representations within the latent space, improving adversarial transferability. Additionally, a perturbation propagation technique is introduced to handle different unknown frame sampling strategies. Experimental results demonstrate that our method can generate adversarial examples that exhibit strong transferability across different V-MLLMs on multiple video-text multimodal tasks. Compared to white-box attacks on these models, our black-box attacks (using BLIP-2 as a surrogate model) achieve competitive performance, with average attack success rate (AASR) of 57.98% on MSVD-QA and 58.26% on MSRVTT-QA for Zero-Shot VideoQA tasks, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。