让图像水印在视频生成中仍能被准确识别。
WaTeRFlow: Watermark Temporal Robustness via Flow Consistency
- 用视频扩散模型模拟真实编辑,训练时注入多种失真。
- 通过光流对齐和时序一致性损失提升每帧检测精度。
- 适合需要跨模态版权保护的视频生成应用。
图像水印可保障真实性与来源追溯,但面对各类失真和强大生成编辑仍易被绕过。基于深度学习的水印方案虽提升了对扩散式图像编辑的鲁棒性,但在图像转视频(I2V)过程中,因逐帧检测能力下降而存在漏洞。I2V已从短而抖动的片段发展为多秒、时序连贯的场景,广泛应用于内容创作、世界建模与仿真流程,跨模态水印恢复因此变得至关重要。本文提出WaTeRFlow框架,专为I2V环境下的水印鲁棒性设计:(i) FUSE(流引导统一合成引擎),通过指令驱动编辑和快速视频扩散代理,在训练中引入真实失真;(ii) 光流变形结合时序一致性损失(TCL),稳定逐帧预测;(iii) 语义保持损失,确保条件信号不丢失。在代表性I2V模型上的实验表明,该方法能准确从视频帧中恢复水印,第一帧和逐帧比特准确率更高,且在视频生成前后施加多种失真时仍具强韧性。
原文摘要 · Abstract (English)
Image watermarking supports authenticity and provenance, yet many schemes are still easy to bypass with various distortions and powerful generative edits. Deep learning-based watermarking has improved robustness to diffusion-based image editing, but a gap remains when a watermarked image is converted to video by image-to-video (I2V), in which per-frame watermark detection weakens. I2V has quickly advanced from short, jittery clips to multi-second, temporally coherent scenes, and it now serves not only content creation but also world-modeling and simulation workflows, making cross-modal watermark recovery crucial. We present WaTeRFlow, a framework tailored for robustness under I2V. It consists of (i) FUSE (Flow-guided Unified Synthesis Engine), which exposes the encoder-decoder to realistic distortions via instruction-driven edits and a fast video diffusion proxy during training, (ii) optical-flow warping with a Temporal Consistency Loss (TCL) that stabilizes per-frame predictions, and (iii) a semantic preservation loss that maintains the conditioning signal. Experiments across representative I2V models show accurate watermark recovery from frames, with higher first-frame and per-frame bit accuracy and resilience when various distortions are applied before or after video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。