给图像加隐蔽干扰,让文本驱动视频生成出错
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation

- 通过削弱模型对文本的依赖,强化视觉线索来保护图像
- 在两种主流视频生成模型上均有效防止内容被非法动画化
- 适合关注版权与隐私的创作者和平台使用
近期图像到视频生成技术可在文本引导下将单张图像转化为逼真视频,带来严重的版权与隐私风险。本文提出 Anti-Prompt,一种图像保护方法:向图像中注入人眼无法察觉的扰动,使文本引导下的图像转视频生成出现明显不一致和结构失败。该方法基于一个简单观察:当移除文本引导时,现代图像到视频模型的生成质量显著下降,不仅运动真实感降低,且主体保留、结构连贯性和时间一致性也受损。基于此,Anti-Prompt 在去噪过程中减弱文本条件交互,增强仅依赖视觉的生成路径。为系统评估保护效果,我们引入基于 Video-LLM 的评估协议,提供可解释、逐帧分析生成瑕疵与不一致性的方法。在两种代表性 I2V 架构上的实验表明,本方法在保持高效性的同时,具备强保护能力与跨模型迁移性。
原文摘要 · Abstract (English)
Recent advances in Image-to-Video generation allow a single image to be animated into a convincing video under text guidance, raising serious copyright and privacy risks. We propose Anti-Prompt, an image protection approach that injects imperceptible perturbations into an image, inducing visible inconsistencies and structural failures in text-guided I2V generation. Our method is motivated by a simple empirical observation. When text guidance is removed from modern I2V models, generation quality degrades markedly, not only in motion realism but also in subject preservation, structural coherence, and temporal consistency. Building on this insight, Anti-Prompt exploits the model reliance on textual guidance by attenuating text-conditioned interactions during denoising while strengthening visual-only pathways. To further systematically evaluate protection effectiveness, we introduce a Video-LLM-assisted evaluation protocol that provides interpretable, frame-grounded analyses of generation artifacts and inconsistencies. Experiments on two representative I2V architectures demonstrate that our method achieves strong protection performance while improving efficiency and cross-model transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。