arXiv:2509.22292cs.CVcs.AI2025-09被引 8

通过拆分场景让文本生成视频模型输出有害内容,突破安全防护。

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

论文配图:Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
图 1 · 摘自论文原文
  • 将恶意叙述拆成多个无害小场景,组合后诱导生成有害视频。
  • 在5个主流模型上攻击成功率最高达84.1%,显著超越现有方法。
  • 适合研究视频生成安全或对抗攻击的学者与工程师参考。

随着文本生成视频(T2V)模型的快速发展,其安全风险日益突出。尽管已有研究针对大语言模型、视觉语言模型及文本生成图像模型开展越狱攻击,但对T2V模型的安全漏洞仍缺乏系统探索。为此,本文提出SceneSplit——一种新型黑盒越狱方法,通过将有害叙事拆分为多个单独看似无害的场景,利用这些场景的序列组合在生成空间中施加强约束,从而引导模型输出有害视频。单个场景对应宽泛且安全的输出空间,但多场景组合后会显著缩小该空间,使其落入不安全区域,极大提升生成有害视频的概率。该机制通过迭代场景调整进一步绕过安全过滤器,并结合可复用的攻击模式库提升攻击效果与鲁棒性。在涵盖11类安全风险的T2VSafetyBench测试中,SceneSplit在Luma Ray2、Hailuo、Veo2、Kling V1.0和Sora2上分别取得77.2%、84.1%、78.2%、78.6%和68.6%的平均攻击成功率,显著优于现有基线。本工作揭示当前T2V模型安全机制在叙事结构层面存在薄弱环节,为理解与改进T2V安全性提供了新视角。

原文摘要 · Abstract (English)

Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While recent studies have explored vulnerabilities in models like LLMs, VLMs, and Text-to-Image (T2I) models through jailbreak attacks, T2V models remain largely unexplored, leaving a significant safety gap. To address this gap, we introduce SceneSplit, a novel black-box jailbreak method that works by fragmenting a harmful narrative into multiple scenes, each individually benign. This approach manipulates the generative output space, the abstract set of all potential video outputs for a given prompt, using the combination of scenes as a powerful constraint to guide the final outcome. While each scene individually corresponds to a wide and safe space where most outcomes are benign, their sequential combination collectively restricts this space, narrowing it to an unsafe region and significantly increasing the likelihood of generating a harmful video. This core mechanism is further enhanced through iterative scene manipulation, which bypasses the safety filter within this constrained unsafe region. Additionally, a strategy library that reuses successful attack patterns further improves the attack's overall effectiveness and robustness. To validate our method, we evaluate SceneSplit across 11 safety categories from T2VSafetyBench on T2V models. Our results show that it achieves a high average Attack Success Rate (ASR) of 77.2% on Luma Ray2, 84.1% on Hailuo, 78.2% on Veo2, 78.6% on Kling V1.0, and 68.6% on Sora2, significantly outperforming the existing baselines. Through this work, we demonstrate that current T2V safety mechanisms are vulnerable to attacks that exploit narrative structure, providing new insights for understanding and improving the safety of T2V models.

文本生成视频越狱攻击安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。