arXiv:2606.02111cs.CVcs.AI2026-06ACL

用多片段视频测试并揭示MLLM安全漏洞的根源

Jailbreaking Multimodal Large Language Models using Multi-Clip Video

论文配图:Jailbreaking Multimodal Large Language Models using Multi-Clip Video
图 1 · 摘自论文原文
  • 设计2920个多片段视频,研究视频多样性对攻击效果的影响
  • 发现视频中片段越多、动态越强、场景越多样,攻击成功率越高
  • 提出利用图像模态更强鲁棒性的防御策略,适合安全研究人员参考

随着多模态大语言模型(MLLMs)发展出处理视频输入的能力,其潜在的恶意滥用风险日益受到关注。已有研究表明,通过视觉输入可绕过MLLM的安全对齐机制,但视频输入中哪些属性引发这种漏洞尚不明确。为此,我们构建了包含2,920个视频的Multi-Clip Video(MCV)SafetyBench数据集,每个视频由多个短片段组成,涵盖与有害查询相关的多样化情境。在8个代表性视频MLLM上的实验表明,攻击成功率随片段数量增加而持续上升。结果进一步显示:(1)视频模态比图像模态更易受攻击;(2)动态视频比静态视频更易被攻破;(3)包含更多元情境的视频更具攻击性。基于此,我们提出一种防御策略,利用图像模态相对更强的鲁棒性来增强系统安全性。

原文摘要 · Abstract (English)

As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. Prior jailbreak studies have shown that safety alignment in MLLMs can be bypassed through visual inputs, yet it remains unclear which properties of video inputs induce this vulnerability. To address this gap, we introduce Multi-Clip Video (MCV) SafetyBench, a dataset of 2,920 videos designed to evaluate how the diversity of video inputs affects the vulnerability of MLLMs. Each video consists of multiple short clips depicting diverse contexts related to a harmful query. Experiments on eight representative video MLLMs show that attack success consistently increases with the number of clips. Our results further indicate that the video modality is (1) more vulnerable than the image modality, (2) more vulnerable to dynamic videos than to static videos, and (3) more vulnerable when videos contain more diverse contexts. Building on these findings, we propose a defense strategy that leverages the relative robustness of the image modality.

多模态安全视频生成对抗攻击防御策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。