提出免疫图像对抗视频生成,防止静态图被恶意动画化。
Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
- 在编码器层强制时序平衡的潜在差异,避免噪声被稀释。
- 通过预计算崩溃轨迹对齐生成过程,抵抗文本引导干扰。
- 在相同隐蔽性下,效果远超传统图像防御方法。
图像到视频(I2V)生成可能带来社会危害,因可未经授权将静态图像转化为逼真的深度伪造视频。现有防御虽能抵御静态图像篡改,但扩展至I2V仍不充分且复杂。本文系统分析现代I2V模型为何对简单图像级对抗攻击(即免疫)具有高度鲁棒性:视频编码过程迅速稀释未来帧中的对抗噪声,而持续的文本条件引导主动覆盖免疫意图。基于此,我们提出Immune2V框架,在编码器层级强制时序平衡的潜在差异以防止信号稀释,并将中间生成表示与预计算的崩溃诱导轨迹对齐,以对抗文本引导的覆盖。大量实验表明,相较于适配的图像级基线,Immune2V在相同不可察觉预算下产生更显著、更持久的退化效果。
原文摘要 · Abstract (English)
Image-to-video (I2V) generation has the potential for societal harm because it enables the unauthorized animation of static images to create realistic deepfakes. While existing defenses effectively protect against static image manipulation, extending these to I2V generation remains underexplored and non-trivial. In this paper, we systematically analyze why modern I2V models are highly robust against naive image-level adversarial attacks (i.e., immunization). We observe that the video encoding process rapidly dilutes the adversarial noise across future frames, and the continuous text-conditioned guidance actively overrides the intended disruptive effect of the immunization. Building on these findings, we propose the Immune2V framework which enforces temporally balanced latent divergence at the encoder level to prevent signal dilution, and aligns intermediate generative representations with a precomputed collapse-inducing trajectory to counteract the text-guidance override. Extensive experiments demonstrate that Immune2V produces substantially stronger and more persistent degradation than adapted image-level baselines under the same imperceptibility budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。