参考图像虽提升视频生成一致性,却暗藏安全漏洞,可被利用进行越狱攻击。
The Price of Consistency: Exploiting Visual Anchors for Multimodal Jailbreaking in Video Generation
- 将有害意图拆分为静态图像和动态文本,实现无需训练的多模态越狱。
- 在主流视频生成平台测试中,成功率显著高于纯文本攻击方法。
- 揭示一致性机制反噬安全的隐患,适合安全研究与对抗样本检测者关注。
视频生成技术正从纯文本驱动转向多条件可控生成,参考图像作为视觉锚点极大提升了时空一致性。然而,这种一致性机制反而抑制了模型从有害内容向良性内容自然漂移的能力,使有害意图难以被消除,从而加剧安全风险——即一致性的代价。本文揭示此视觉锚定效应,提出无需训练的多模态越狱框架DIVA:将有害意图解耦为静态视觉锚图与动态运动文本提示,并采用双标准选择策略平衡攻击隐蔽性与语义保真度。在多个主流商业平台及开源视频生成模型上的实验表明,DIVA攻击成功率远超现有纯文本方法。为推动后续研究,我们构建首个多条件视频生成安全基准TI2VSafetyBench。
原文摘要 · Abstract (English)
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significantly enhance controllability, their impact on safety remains largely unexplored. In this work, we reveal the visual anchoring effect: by enforcing consistency, the mechanism prevents the generated content from drifting away from the original harmful intent, thereby eliminating the model's natural safety escape route from harmful to benign content. Consequently, visual anchors inherently increase the safety risk---this is the price of consistency. Building on this insight, we propose Decoupling Intent via Visual Anchors (DIVA), a training-free multimodal jailbreak framework for video generation that exploits this vulnerability. DIVA decouples harmful intent into a static visual anchor image and a dynamic motion text prompt, and employs dual-criteria selection to balance attack stealthiness with semantic preservation. Extensive experiments across various leading commercial platforms and mainstream open-source video generation models demonstrate that DIVA achieves a substantially higher Attack Success Rate than existing text-only methods. To facilitate future research, we additionally contribute TI2VSafetyBench, the first safety benchmark for multi-conditional video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。