通过只给视频首尾帧,让文本生成视频模型自动补全有害内容。
Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking
- 用仅包含首尾帧的碎片化提示,诱导模型自主生成中间有害画面。
- 在商用模型上攻击成功率提升12%,远超传统改写方法。
- 揭示了视频生成模型对时间轨迹的隐式补全风险,适合安全研究者关注。
近期文本到视频(T2V)模型能从简短自然语言提示生成复杂视频,引发真实世界滥用的安全隐患。以往越狱攻击多通过改写提示避开内容过滤,但往往仍保留明显敏感线索,忽视了视频生成特有的深层缺陷。本文发现:当提示仅指定稀疏边界条件(如起始与结束帧),而中间演化过程未明确时,模型可能自主重建包含有害内容的合理时间轨迹,尽管输入输出均看似无害。基于此,提出TFM框架,将原不安全请求转化为仅含两帧的时间片段提取,并通过隐式替换进一步隐藏敏感线索。在多个开源及商用T2V模型上的评估显示,TFM显著提升越狱成功率,商业系统最高达12%的提升。结果表明,需建立考虑模型自主补全能力的时序感知安全机制。
原文摘要 · Abstract (English)
Recent text-to-video (T2V) models can synthesize complex videos from lightweight natural language prompts, raising urgent concerns about safety alignment in the event of misuse in the real world. Prior jailbreak attacks typically rewrite unsafe prompts into paraphrases that evade content filters while preserving meaning. Yet, these approaches often still retain explicit sensitive cues in the input text and therefore overlook a more profound, video-specific weakness. In this paper, we identify a temporal trajectory infilling vulnerability of T2V systems under fragmented prompts: when the prompt specifies only sparse boundary conditions (e.g., start and end frames) and leaves the intermediate evolution underspecified, the model may autonomously reconstruct a plausible trajectory that includes harmful intermediate frames, despite the prompt appearing benign to input or output side filtering. Building on this observation, we propose TFM. This fragmented prompting framework converts an originally unsafe request into a temporally sparse two-frame extraction and further reduces overtly sensitive cues via implicit substitution. Extensive evaluations across multiple open-source and commercial T2V models demonstrate that TFM consistently enhances jailbreak effectiveness, achieving up to a 12% increase in attack success rate on commercial systems. Our findings highlight the need for temporally aware safety mechanisms that account for model-driven completion beyond prompt surface form.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。