arXiv:2605.01761cs.CV2026-05

提出TrajShield防御文本生成视频的越狱攻击与隐性安全风险。

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

论文配图:TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks
图 1 · 摘自论文原文
  • 通过模拟提示语的语义演化轨迹定位风险源头
  • 在14类安全风险上平均降低52.44%攻击成功率
  • 无需训练,适用于多种视频生成模型

文本生成视频(T2V)模型虽能生成连贯视频,但存在生成暴力或色情内容等安全风险。现有基于词法层面的防护方法易受重述或对抗性提示的越狱攻击。此外,T2V还面临一种独特挑战:时序涌现风险——看似无害的提示在生成过程中因叙事连贯性推演而产生不安全内容。本文提出TrajShield,一种无需训练、在推理阶段运行的安全防御框架,将T2V安全问题建模为时序语义空间中的因果干预。该方法通过模拟提示的潜在演化轨迹,定位风险根源,并施加最小化语义扰动以消除风险,同时保留无关安全性的内容。在涵盖14类安全风险的T2VSafetyBench测试中,对多个T2V模型验证显示,TrajShield实现领先防御效果,平均攻击成功率降低52.44%,且保持高语义保真度。

原文摘要 · Abstract (English)

Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also risk producing unsafe content such as violence or explicit material. Existing prompt-level defenses are largely inherited from text-to-image safety and operate on the lexical surface of the input, making them vulnerable to jailbreak attacks that disguise harmful intent through rephrasing or adversarial prompting. Moreover, T2V generation introduces a distinctive challenge overlooked by prior work: temporally emergent risk, where a seemingly benign prompt leads to unsafe content through the generator's temporal extrapolation toward narrative coherence. We propose \method{}, a training-free, inference-time defense framework that reformulates T2V safety as a causal intervention in a temporally structured semantic space. TrajShield handles explicit unsafe prompts, jailbreak attacks, and temporally emergent risks in a unified manner by simulating the implied trajectory of a prompt, localizing the causal origin of potential risk, and applying a minimally invasive rewrite that neutralizes the risk while preserving safety-irrelevant semantics. Experiments on T2VSafetyBench across 14 safety categories and multiple T2V backends demonstrate that TrajShield achieves state-of-the-art defenseive performance while maintaining high semantic fidelity, substantially outperforming existing defenses, with an average ASR reduction of 52.44\%.

视频生成安全防御越狱攻击时序风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。