用隐晦提示让文本生成视频模型违规输出,攻击更隐蔽有效。
SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge
- 设计三模块提示:场景锚点保真实,声音线索引异常,风格指令强化效果。
- 在7个主流模型上平均攻击成功率提升23%,商业模型效果显著。
- 适合研究模型安全、对抗攻击的学者和工程师参考。
越狱攻击可绕过模型安全机制,暴露关键漏洞。以往针对文本到视频(T2V)模型的攻击多在明显不安全提示中添加对抗扰动,易被检测和防御。本文提出SPARK框架,利用T2V模型跨模态关联特性,通过模块化提示实现隐蔽攻击。提示包含三部分:中性场景锚点(从被屏蔽意图中提取表面场景描述,维持合理性)、潜在声音触发器(如吱呀声、闷响等看似无害的音频描述,利用学习到的音视频共现先验诱导不安全视觉概念)、风格调制器(如镜头构图、氛围设定,放大并稳定触发器效果)。将攻击生成建模为该模块化提示空间上的约束优化问题,并采用引导搜索算法平衡隐蔽性与有效性。在7个T2V模型上的实验表明,该方法在商业模型上平均攻击成功率提升23%。
原文摘要 · Abstract (English)
Jailbreak attacks can circumvent model safety guardrails and reveal critical blind spots. Prior attacks on text-to-video (T2V) models typically add adversarial perturbations to obviously unsafe prompts, which are often easy to detect and defend. In contrast, we show that benign-looking prompts containing rich, implicit cues can induce T2V models to generate semantically unsafe videos that both violate policy and preserve the original (blocked) intent. To realize this, we propose SPARK, a jailbreak framework that leverages T2V models cross-modal associative patterns via a modular prompt design. Specifically, our prompts combine three components: neutral scene anchors, which provide the surface-level scene description extracted from the blocked intent to maintain plausibility; latent auditory triggers, textual descriptions of innocuous-sounding audio events (e.g., creaking, muffled noises) that exploit learned audio-visual co-occurrence priors to bias the model toward particular unsafe visual concepts; and stylistic modulators, cinematic directives (e.g., camera framing, atmosphere) that amplify and stabilize the latent trigger's effect. We formalize attack generation as a constrained optimization over the above modular prompt space and solve it with a guided search procedure that balances stealth and effectiveness. Extensive experiments over 7 T2V models demonstrate the efficacy of our attack, achieving a +23% improvement in average attack success rate in commercial models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。