arXiv:2608.19737cs.CVcs.AI2026-08

通过精心调度字幕时间,突破大模型视频安全防护。

TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling

论文配图:TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
图 1 · 摘自论文原文
  • 用可精准控制时间的字幕构造对话式攻击序列
  • 在多个模型上实现超50%的攻击成功率,最高提升53个百分点
  • 适合研究视频安全、对抗攻击或模型鲁棒性的学者

大型视觉语言模型(LVLMs)在视频理解与推理方面取得了显著进展。尽管已有大量关于文本和图像的越狱攻击研究,针对LVLMs的视频越狱仍基本未被探索。现有方法主要操控视频中的文本内容,却忽视了信息的时间组织方式。我们的分析表明,越狱效果不仅取决于文本语义,还受其呈现时间(如持续时长和时间段分配)影响。基于此,我们利用真实视频中常见的字幕作为攻击媒介——其既能精确控制语义呈现时间,又不易被视觉察觉。为此,提出黑盒视频越狱框架TempJail,构建与查询对齐的对话式字幕序列,并优化其时间调度,以挖掘LVLMs的时间漏洞,诱导产生符合原始恶意意图的响应。在四个代表性LVLMs和两个数据集上的实验表明,TempJail在所有评估设置中均达到最高攻击成功率,相较于最强基线,在GPT-5和Gemini 3.5-Flash的数据集平均攻击成功率(ASR)上分别提升53和18个百分点。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored. Existing video jailbreak methods mainly manipulate textual content embedded in videos, while overlooking how such information is organized over time. Our analysis reveals that jailbreak effectiveness depends not only on the semantics of textual information but also on its temporal presentation, including duration and timing-slot allocation. Motivated by this finding, we use subtitles, which are common in real-world videos and allow semantic content to be presented under precise temporal control without appearing visually intrusive, as a natural attack medium. Based on this insight, we propose TempJail, a black-box video-based jailbreak framework that constructs query-aligned dialogue-style subtitle sequences and optimizes their temporal scheduling to exploit temporal vulnerabilities in LVLMs and elicit responses that satisfy the harmful intent of the source query. Extensive experiments on four representative LVLMs and two datasets demonstrate that TempJail achieves the highest attack success rate across all evaluated model--dataset settings, outperforming the strongest baseline by 53 and 18 percentage points in dataset-averaged ASR on GPT-5 and Gemini 3.5-Flash, respectively.

视频越狱对抗攻击视觉语言模型字幕调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。