arXiv:2512.16874cs.CVcs.AI2025-12被引 7

提出无需像素损失的对抗训练法,实现图像视频水印的隐形与强鲁棒性。

Pixel Seal: Adversarial-only training for invisible image and video watermarking

  • 纯对抗训练,摒弃易失真的像素损失,更贴近人眼感知。
  • 三阶段训练稳定收敛,高分辨率下仍保持水印不可见且抗干扰。
  • 支持视频场景,可高效适配真实应用,适合内容溯源需求者。

隐形水印对数字内容溯源至关重要,但现有方法常难以平衡鲁棒性与不可见性。本文提出Pixel Seal,针对三大问题:(i)依赖MSE、LPIPS等代理感知损失,导致可见水印痕迹;(ii)目标冲突引发优化不稳,需大量调参;(iii)模型扩展至高分辨率时鲁棒性与不可见性下降。为此,提出纯对抗训练范式,移除不可靠的像素级损失;设计三阶段训练流程,解耦鲁棒性与不可见性以稳定收敛;通过基于JND的衰减和训练时推理模拟,解决高分辨率缩放伪影。在多种图像类型及广泛变换下评估显示,性能显著优于现有方法。最后证明其可通过时间水印池化高效适配视频,成为真实场景中可靠溯源的实用方案。

原文摘要 · Abstract (English)

Invisible watermarking is essential for tracing the provenance of digital content. However, training state-of-the-art models remains notoriously difficult, with current approaches often struggling to balance robustness against true imperceptibility. This work introduces Pixel Seal, which sets a new state-of-the-art for image and video watermarking. We first identify three fundamental issues of existing methods: (i) the reliance on proxy perceptual losses such as MSE and LPIPS that fail to mimic human perception and result in visible watermark artifacts; (ii) the optimization instability caused by conflicting objectives, which necessitates exhaustive hyperparameter tuning; and (iii) reduced robustness and imperceptibility of watermarks when scaling models to high-resolution images and videos. To overcome these issues, we first propose an adversarial-only training paradigm that eliminates unreliable pixel-wise imperceptibility losses. Second, we introduce a three-stage training schedule that stabilizes convergence by decoupling robustness and imperceptibility. Third, we address the resolution gap via high-resolution adaptation, employing JND-based attenuation and training-time inference simulation to eliminate upscaling artifacts. We thoroughly evaluate the robustness and imperceptibility of Pixel Seal on different image types and across a wide range of transformations, and show clear improvements over the state-of-the-art. We finally demonstrate that the model efficiently adapts to video via temporal watermark pooling, positioning Pixel Seal as a practical and scalable solution for reliable provenance in real-world image and video settings.

水印对抗训练视频生成图像溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。