arXiv:2607.13336cs.CVcs.CR2026-07被引 1

提出首个统一防御图像转视频与微调定制的视频隐私保护方法

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

论文配图:Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization
图 1 · 摘自论文原文
  • 通过滑动窗口多帧扰动,融合时序压缩特性实现跨视频通用防护
  • 在参考式与微调式攻击下均显著提升身份保护效果,对未见时序攻击鲁棒
  • 适合关注视频生成隐私安全的研究者与应用开发者

基于扩散模型的视频生成技术可通过微调或参考图像实现高质量个性化定制,但带来隐私泄露、身份归属与知识产权风险。现有防护方法主要针对图像,对视频在参考式与微调式定制下的双重威胁缺乏有效应对。本文提出时序一致的通用对抗扰动(TC-UAP),首次实现对两类视频定制攻击的统一防护。TC-UAP在多个视频的滑动窗口上优化身份级多帧扰动,考虑视频VAE中的时序压缩依赖性,使单一扰动可保护不同长度的未见视频。引入内在时序建模与外在代理时序攻击损失,增强扰动时序一致性与对未知时序攻击的鲁棒性。实验表明,相比现有方法,TC-UAP在参考式与微调式攻击下均实现更强的身份保护,并在多种未见时序攻击中保持稳定性能。

原文摘要 · Abstract (English)

Recent diffusion-based video generation models have enabled high-quality personalized video customization through both tuning-based pipelines, which fine-tune a video diffusion model, and reference-based pipelines such as image-to-video generation. However, these capabilities raise serious concerns about personal privacy, identity ownership and intellectual property protection. Existing anti-customization works focus on protecting images, while protection for videos against both reference- and tuning-based customization remains largely underexplored. Protecting videos in this setting raises three challenges: (i) Image-level perturbations, optimized frame by frame, cannot survive temporal compression by 3D video VAE. (ii) A video-level perturbation optimized on a single video is vulnerable to temporal editing and fails to protect unseen videos. (iii) Temporally inconsistent perturbations are not robust to temporal attacks. To address these challenges, we propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization. TC-UAP optimizes an identity-level multi-frame UAP over sliding windows from multiple videos, accounting for local temporal dependencies induced by temporal compression in video VAE and enabling a single perturbation to protect unseen videos of varying lengths. Moreover, we introduce intrinsic temporal modeling and an extrinsic surrogate temporal-attack loss, which make the perturbation temporally consistent and robust to unseen temporal attacks. Empirically, quantitative and qualitative results show that TC-UAP achieves the strongest identity protection compared with existing methods under both reference- and tuning-based video customization, and remains robust under multiple unseen temporal attacks.

视频生成隐私保护对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。