arXiv:2602.20999cs.CV2026-02被引 9

用图像伪装恶意指令,突破图像生成模型安全限制。

VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models

  • 将危险文本意图转为图像中的隐蔽视觉指令。
  • 在4个主流模型上实现最高83.5%攻击成功率,拒答率接近零。
  • 无需训练、可迁移,适合研究模型安全与对抗攻击者。

图像到视频(I2V)生成模型能基于参考图像生成视频,展现出一定的视觉指令跟随能力,使图像中的视觉线索可作为隐式控制信号。然而,这种能力也带来了未被重视的风险:攻击者可通过图像模态注入恶意意图。本文提出视觉指令注入(VII),一种无需训练且可迁移的越狱框架,将不安全文本提示中的恶意意图伪装成安全参考图像中的良性视觉指令。VII通过恶意意图重编程模块从不安全文本中提取恶意意图,同时最小化其显性危害,并利用视觉指令定位模块将该意图以语义一致的方式嵌入安全图像中,从而在视频生成过程中诱导有害内容输出。在四个先进商业I2V模型(Kling-v2.5-turbo、Gemini Veo-3.1、Seedance-1.5-pro、PixVerse-V5)上的实验表明,VII实现了最高83.5%的攻击成功率,同时将拒答率降至接近零,显著优于现有基线方法。

原文摘要 · Abstract (English)

Image-to-Video (I2V) generation models, which condition video generation on reference images, have shown emerging visual instruction-following capability, allowing certain visual cues in reference images to act as implicit control signals for video generation. However, this capability also introduces a previously overlooked risk: adversaries may exploit visual instructions to inject malicious intent through the image modality. In this work, we uncover this risk by proposing Visual Instruction Injection (VII), a training-free and transferable jailbreaking framework that intentionally disguises the malicious intent of unsafe text prompts as benign visual instructions in the safe reference image. Specifically, VII coordinates a Malicious Intent Reprogramming module to distill malicious intent from unsafe text prompts while minimizing their static harmfulness, and a Visual Instruction Grounding module to ground the distilled intent onto a safe input image by rendering visual instructions that preserve semantic consistency with the original unsafe text prompt, thereby inducing harmful content during I2V generation. Empirically, our extensive experiments on four state-of-the-art commercial I2V models (Kling-v2.5-turbo, Gemini Veo-3.1, Seedance-1.5-pro, and PixVerse-V5) demonstrate that VII achieves Attack Success Rates of up to 83.5% while reducing Refusal Rates to near zero, significantly outperforming existing baselines.

图像生成越狱攻击安全风险视觉指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。