I2VShield用轻量攻击保护图像转视频不被滥用
I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

- 用自适应扰动生成降低计算开销,保持视觉不可见
- 通过多模态注意力破坏让视频时空连贯性崩溃
- 适合资源有限的防御方使用,针对性强
图像转视频(I2V)模型的快速发展带来了滥用风险。尽管已有大量研究聚焦于检测生成视频,但对I2V模型的主动防御仍缺乏探索。现有方法多依赖基于梯度的对抗攻击,需大量显存(VRAM)生成对抗样本。为此,本文提出I2VShield,一种针对扩散变压器(DiT)架构的I2V模型的隐私保护方法。该方法包含两部分:(1) 文本自适应扰动生成框架,结合对抗学习,在保持视觉不可察觉的同时降低计算开销;(2) 无目标多模态注意力破坏(MAD)攻击,利用DiT模型内在漏洞,最大化内部注意力特征偏离原始状态。大量实验表明,该方法在多个数据集和主流DiT-based I2V模型上均表现出优异的防护性能,尤其能有效破坏视频的时空连贯性,同时显著降低计算成本。
原文摘要 · Abstract (English)
The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has been made in detecting AI-generated videos, proactive defenses against I2V models remain underexplored. In particular, current proactive defenses against I2V models predominantly rely on gradient-based adversarial attacks, which require defenders to possess GPUs with substantial memory resources (VRAM) to generate adversarial examples. To address this issue, we propose I2VShield, a privacy protection method based on generative adversarial attacks tailored to Diffusion Transformer (DiT)-based I2V models. The proposed method primarily consists of two components: (1) a text-adaptive perturbation generation framework integrating adversarial learning to mitigate computational overhead while maintaining visual imperceptibility; and (2) an untargeted Multimodal Attention Disruption (MAD) attack that exploits the inherent vulnerabilities of DiT-based I2V models, maximizing the deviation of the internal attention features from their clean states. Extensive experiments demonstrate that our approach achieves highly competitive protection performance across various datasets and mainstream DiT-based I2V models, particularly in disrupting spatiotemporal coherence, while substantially reducing computational costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。