提出离散优化方法,让文本生成视频模型更易被诱导输出违规内容。
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
- 将文本到视频越狱攻击建模为离散优化问题,联合优化逃逸率与内容一致性。
- 在多个开源和闭源模型上测试,攻破成功率比现有方法提升11.4%和10.0%。
- 适合安全研究者评估模型漏洞,或用于对抗性防御训练。
近年来,随着扩散模型的快速发展,文本到视频(T2V)生成模型取得了显著进展,代表性模型包括Pika、Luma、Kling和Open-Sora。尽管这些模型具备强大的生成能力,但其对越狱攻击的脆弱性也暴露了重大安全风险,即模型可能被操控生成色情、暴力或歧视性内容。现有工作如T2VSafetyBench提供了初步的安全评估基准,但缺乏系统性的漏洞探索方法。为此,我们首次将T2V越狱攻击形式化为离散优化问题,并提出基于联合目标的优化框架T2V-OptJail。该框架包含两个关键优化目标:绕过内置安全过滤机制以提高攻击成功率,同时保持对抗提示与恶意输入提示之间、生成视频与恶意输入提示之间的语义一致性,以增强内容可控性。此外,引入基于提示变体的迭代优化策略,每轮生成多个语义等价候选提示,通过聚合得分稳健引导搜索至最优对抗提示。我们在多个T2V模型上进行了大规模实验,涵盖开源与真实商业闭源模型。结果表明,该方法在GPT-4评估的攻击成功率上提升11.4%,在人工评估的攻击成功率上提升10.0%,验证了其在攻击效果与内容控制方面的显著优势。
原文摘要 · Abstract (English)
In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit impressive generative capabilities, they also expose significant security risks due to their vulnerability to jailbreak attacks, where the models are manipulated to produce unsafe content such as pornography, violence, or discrimination. Existing works such as T2VSafetyBench provide preliminary benchmarks for safety evaluation, but lack systematic methods for thoroughly exploring model vulnerabilities. To address this gap, we are the first to formalize the T2V jailbreak attack as a discrete optimization problem and propose a joint objective-based optimization framework, called T2V-OptJail. This framework consists of two key optimization goals: bypassing the built-in safety filtering mechanisms to increase the attack success rate, preserving semantic consistency between the adversarial prompt and the unsafe input prompt, as well as between the generated video and the unsafe input prompt, to enhance content controllability. In addition, we introduce an iterative optimization strategy guided by prompt variants, where multiple semantically equivalent candidates are generated in each round, and their scores are aggregated to robustly guide the search toward optimal adversarial prompts. We conduct large-scale experiments on several T2V models, covering both open-source models and real commercial closed-source models. The experimental results show that the proposed method improves 11.4% and 10.0% over the existing state-of-the-art method in terms of attack success rate assessed by GPT-4, attack success rate assessed by human accessors, respectively, verifying the significant advantages of the method in terms of attack effectiveness and content control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。