arXiv:2603.24570cs.CVcs.AI2026-03中稿 · CVPR被引 6

对抗用照片生成假视频的恶意攻击,保护个人隐私

Anti-I2V: Safeguarding your photos from malicious image-to-video generation

  • 在颜色与频域双重空间加扰动,增强防御鲁棒性
  • 针对扩散模型关键层设计训练目标,显著破坏视频连贯性
  • 适用于多种架构,尤其对DiT类模型有效

基于扩散模型的图像到视频生成技术虽提升了人物动画质量,却可能被滥用于从特定人物照片和文本提示生成伪造视频。现有防御多针对图像生成,较少关注图像到视频扩散模型(VDMs),且主要聚焦于UNet架构,对具备更强特征保留与时间一致性的扩散变压器(DiT)模型的防御效果尚不明确。本文提出Anti-I2V,一种适用于多种扩散骨干网络的新型防御机制。该方法不仅在RGB空间,还在$L$*$a$*$b$*颜色空间和频率域引入扰动,集中作用于显著像素。进一步识别去噪过程中捕捉最显著语义特征的网络层,设计针对性训练目标以最大化降低生成视频的时间连贯性与保真度。大量实验表明,Anti-I2V在抵御多种视频扩散模型方面达到当前最优防御性能,为应对恶意生成提供了有效解决方案。

原文摘要 · Abstract (English)

Advances in diffusion-based video generation models, while significantly improving human animation, poses threats of misuse through the creation of fake videos from a specific person's photo and text prompts. Recent efforts have focused on adversarial attacks that introduce crafted perturbations to protect images from diffusion-based models. However, most existing approaches target image generation, while relatively few explicitly address image-to-video diffusion models (VDMs), and most primarily focus on UNet-based architectures. Hence, their effectiveness against Diffusion Transformer (DiT) models remains largely under-explored, as these models demonstrate improved feature retention, and stronger temporal consistency due to larger capacity and advanced attention mechanisms. In this work, we introduce Anti-I2V, a novel defense against malicious human image-to-video generation, applicable across diverse diffusion backbones. Instead of restricting noise updates to the RGB space, Anti-I2V operates in both the $L$*$a$*$b$* and frequency domains, improving robustness and concentrating on salient pixels. We then identify the network layers that capture the most distinct semantic features during the denoising process to design appropriate training objectives that maximize degradation of temporal coherence and generation fidelity. Through extensive validation, Anti-I2V demonstrates state-of-the-art defense performance against diverse video diffusion models, offering an effective solution to the problem.

视频生成隐私保护扩散模型防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。