T2VShield可无须模型参数,防御文本生成视频的越狱攻击。
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video Models

- 通过推理与多模态检索重写提示词,净化恶意输入。
- 在五个平台测试中,使越狱成功率降低35%。
- 适合构建安全视频模拟器的开发者使用。
生成式人工智能的快速发展使文本到视频模型成为未来多模态世界模拟器的关键。然而,这些模型仍易受越狱攻击,即精心设计的提示绕过安全机制,生成有害内容,威胁模拟应用的可靠性与安全性。本文提出T2VShield,一种全面且模型无关的防御框架,用于保护文本到视频模型免受越狱威胁。该方法系统分析输入、模型和输出阶段,识别现有防御的局限性,包括提示语义模糊、动态视频输出中恶意内容检测困难,以及以模型为中心的应对策略僵化。T2VShield引入基于推理与多模态检索的提示重写机制,以净化恶意输入,并设计多范围检测模块,捕捉时空与跨模态的局部与全局不一致。该框架无需访问内部模型参数,适用于开源与闭源系统。在五个平台上的大量实验表明,相比强基线,其可将越狱成功率降低高达35%。我们还开发了以人为中心的音视频评估协议,评估感知安全性,强调视觉层面防御对提升下一代多模态模拟器可信度的重要性。
原文摘要 · Abstract (English)
The rapid development of generative artificial intelligence has made text to video models essential for building future multimodal world simulators. However, these models remain vulnerable to jailbreak attacks, where specially crafted prompts bypass safety mechanisms and lead to the generation of harmful or unsafe content. Such vulnerabilities undermine the reliability and security of simulation based applications. In this paper, we propose T2VShield, a comprehensive and model agnostic defense framework designed to protect text to video models from jailbreak threats. Our method systematically analyzes the input, model, and output stages to identify the limitations of existing defenses, including semantic ambiguities in prompts, difficulties in detecting malicious content in dynamic video outputs, and inflexible model centric mitigation strategies. T2VShield introduces a prompt rewriting mechanism based on reasoning and multimodal retrieval to sanitize malicious inputs, along with a multi scope detection module that captures local and global inconsistencies across time and modalities. The framework does not require access to internal model parameters and works with both open and closed source systems. Extensive experiments on five platforms show that T2VShield can reduce jailbreak success rates by up to 35 percent compared to strong baselines. We further develop a human centered audiovisual evaluation protocol to assess perceptual safety, emphasizing the importance of visual level defense in enhancing the trustworthiness of next generation multimodal simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。