给视频生成模型嵌入不可见水印,生成无延迟且质量高。
Video Signature: Implicit Watermarking for Video Diffusion Models
- 通过微调潜空间解码器,让水印在生成时自然融入。
- 水印提取准确率高,视频质量与原生生成几乎无差别。
- 适合需要版权保护的AI视频应用,如内容溯源和防伪。
人工智能生成内容(AIGC)在视频生成领域快速发展,但引发了知识产权保护与内容追溯的严峻问题。水印是常见解决方案,但现有方法多为生成后添加,难以兼顾画质与水印可提取性;而生成中嵌入水印的方法通常带来显著计算开销。为此,我们提出 extbf{Video Signature}( extsc{VidSig}),一种针对视频扩散模型的隐式水印方法,可在生成过程中实现几乎无额外延迟的不可见、自适应水印融合。具体而言,我们部分微调潜空间解码器,引入 extbf{感知敏感层抑制}(PAS)机制,预先识别并冻结视觉敏感层以保障画质。为进一步提升时序一致性,我们设计轻量级 extbf{时序对齐}模块,在微调阶段引导解码器生成连贯帧序列。实验表明, extsc{VidSig} 在水印提取准确率、视频质量与水印延迟之间达到最佳平衡,对空间与时间篡改均具强鲁棒性,且在不同视频长度与分辨率下保持稳定,具备实际应用价值。
原文摘要 · Abstract (English)
The rapid development of Artificial Intelligence Generated Content (AIGC) has led to significant progress in video generation, but also raises serious concerns about intellectual property protection and reliable content tracing. Watermarking is a widely adopted solution to this issue, yet existing methods for video generation mainly follow a post-generation paradigm, which often fails to effectively balance the trade-off between video quality and watermark extraction. Meanwhile, current in-generation methods that embed the watermark into the initial Gaussian noise usually incur substantial additional computation. To address these issues, we propose \textbf{Video Signature} (\textsc{VidSig}), an implicit watermarking method for video diffusion models that enables imperceptible and adaptive watermark integration during video generation with almost no extra latency. Specifically, we partially fine-tune the latent decoder, where \textbf{Perturbation-Aware Suppression} (PAS) pre-identifies and freezes perceptually sensitive layers to preserve visual quality. Beyond spatial fidelity, we further enhance temporal consistency by introducing a lightweight \textbf{Temporal Alignment} module that guides the decoder to generate coherent frame sequences during fine-tuning. Experimental results show that \textsc{VidSig} achieves the best trade-off among watermark extraction accuracy, video quality, and watermark latency. It also demonstrates strong robustness against both spatial and temporal tamper, and remains stable across different video lengths and resolutions, highlighting its practicality in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。