arXiv:2512.19048cs.CV2025-12

让图像水印在视频生成中仍能被准确识别。

WaTeRFlow: Watermark Temporal Robustness via Flow Consistency

  • 用视频扩散模型模拟真实编辑,训练时注入多种失真。
  • 通过光流对齐和时序一致性损失提升每帧检测精度。
  • 适合需要跨模态版权保护的视频生成应用。

图像水印可保障真实性与来源追溯,但面对各类失真和强大生成编辑仍易被绕过。基于深度学习的水印方案虽提升了对扩散式图像编辑的鲁棒性,但在图像转视频(I2V)过程中,因逐帧检测能力下降而存在漏洞。I2V已从短而抖动的片段发展为多秒、时序连贯的场景,广泛应用于内容创作、世界建模与仿真流程,跨模态水印恢复因此变得至关重要。本文提出WaTeRFlow框架,专为I2V环境下的水印鲁棒性设计:(i) FUSE(流引导统一合成引擎),通过指令驱动编辑和快速视频扩散代理,在训练中引入真实失真;(ii) 光流变形结合时序一致性损失(TCL),稳定逐帧预测;(iii) 语义保持损失,确保条件信号不丢失。在代表性I2V模型上的实验表明,该方法能准确从视频帧中恢复水印,第一帧和逐帧比特准确率更高,且在视频生成前后施加多种失真时仍具强韧性。

原文摘要 · Abstract (English)

Image watermarking supports authenticity and provenance, yet many schemes are still easy to bypass with various distortions and powerful generative edits. Deep learning-based watermarking has improved robustness to diffusion-based image editing, but a gap remains when a watermarked image is converted to video by image-to-video (I2V), in which per-frame watermark detection weakens. I2V has quickly advanced from short, jittery clips to multi-second, temporally coherent scenes, and it now serves not only content creation but also world-modeling and simulation workflows, making cross-modal watermark recovery crucial. We present WaTeRFlow, a framework tailored for robustness under I2V. It consists of (i) FUSE (Flow-guided Unified Synthesis Engine), which exposes the encoder-decoder to realistic distortions via instruction-driven edits and a fast video diffusion proxy during training, (ii) optical-flow warping with a Temporal Consistency Loss (TCL) that stabilizes per-frame predictions, and (iii) a semantic preservation loss that maintains the conditioning signal. Experiments across representative I2V models show accurate watermark recovery from frames, with higher first-frame and per-frame bit accuracy and resilience when various distortions are applied before or after video generation.

图像水印视频生成时序鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。