提出首个视觉提示攻击基准与防御框架,保护图像转视频模型免受隐性指令侵害。
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

- 构建视觉提示攻击分类基准VVA-Bench,系统评估视频生成安全风险。
- 主流模型在攻击下成功率高达100.0%(Wan 2.7)和74.8%(Veo 3.1)。
- VPA-Guard通过自进化检索增强防御,显著降低危害性并保持正常功能。
图像转视频(I2V)生成技术已将静态图像转化为可交互的控制接口,如箭头、草图和表情符号等视觉线索可操控复杂视频动态。然而,这些看似无害的视觉提示可能被模型解读为可执行的时间指令,导致生成有害内容。现有安全评测多聚焦于文本或纯图像的越狱攻击,对隐性视觉提示攻击关注不足。为此,我们提出VVA-Bench——首个系统性评估视觉中心提示攻击下视频生成安全性的基准。在该基准上,实验表明先进模型极易受此类攻击:在Wan 2.7上攻击成功率达100.0%,在Veo 3.1上为74.8%。为缓解风险,我们提出VPA-Guard,一种基于少样本推理的检索增强与自进化防御框架。该方法平均降低攻击成功率44.2%、危害性评分73.4%,同时保持模型对合法用户编辑的有效性。本工作为安全可控的多模态生成提供了严谨评估与有效防御方案。
原文摘要 · Abstract (English)
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos. Despite the severity of this threat, existing safety benchmarks remain predominantly focused on text-based and content-only image-based jailbreaks, leaving implicit visual prompt attacks insufficiently explored. To bridge this gap, we present VVA-Bench, the first systematic benchmark for evaluating video generation safety under categorized vision-centric prompt attacks. Extensive experiments on VVA-Bench demonstrate that state-of-the-art models are highly susceptible to such attacks, with Attack Success Rates (ASR) reaching 100.0\% on Wan 2.7 and 74.8\% on Veo 3.1. To mitigate these risks, we propose VPA-Guard, a retrieval-augmented and self-evolving defense framework. By leveraging few-shot reasoning to identify latent malicious intents, our method reduces the attack ASR by 44.2\% and the harmfulness score by 73.4\% on average, while maintaining the model's utility for legitimate user edits. Our work provides both a rigorous benchmark and an effective defense strategy to advance safe and socially responsible multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。