arXiv:2508.08667cs.CVcs.MM2025-08

提出分阶段优化框架,实现水印不可见、鲁棒且高效。

Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization

  • 分两阶段优化:先对齐分布建共同潜空间,再分离水印表示。
  • 水印提取准确率比现有方法高7.6%,千张图处理仅需1秒。
  • 适合需要高鲁棒性与低延迟的图像版权保护场景。

深度图像水印技术能实现覆盖图像中不可感知的水印嵌入与可靠提取,有效保护图像资产版权。然而,现有方法难以同时满足三个关键要求:(1) 不可见性(水印隐蔽隐藏),(2) 鲁棒性(在多种条件下可靠恢复水印),(3) 广泛适用性(水印过程低延迟)。为此,本文提出分层水印学习(HiWL)框架,采用两阶段优化,使水印模型同时满足上述三要素。第一阶段通过分布对齐学习构建共同潜空间,施加双重约束:(1) 水印图像与原图在视觉上的一致性,(2) 水印潜在表示的信息不变性。该机制有效表征多模态输入——包括二进制水印消息与RGB像素覆盖图像——确保水印不可见性与鲁棒性。第二阶段采用广义水印表示学习,在RGB空间中分离出独特的水印表示。训练完成后,HiWL模型能学习通用水印表示,同时保持广泛适用性。大量实验表明该方法有效性:相比现有方法,水印提取准确率提升7.6%,且处理1000张图像仅需1秒。

原文摘要 · Abstract (English)

Deep image watermarking, which refers to enabling imperceptible watermark embedding and reliable extraction in cover images, has been shown to be effective for copyright protection of image assets. However, existing methods face limitations in simultaneously satisfying three essential criteria for generalizable watermarking: (1) invisibility (imperceptible hiding of watermarks), (2) robustness (reliable watermark recovery under diverse conditions), and (3) broad applicability (low latency in the watermarking process). To address these limitations, we propose a Hierarchical Watermark Learning (HiWL) framework, a two-stage optimization that enables a watermarking model to simultaneously achieve all three criteria. In the first stage, distribution alignment learning is designed to establish a common latent space with two constraints: (1) visual consistency between watermarked and non-watermarked images, and (2) information invariance across watermark latent representations. In this way, multimodal inputs -- including watermark messages (binary codes) and cover images (RGB pixels) -- can be effectively represented, ensuring both the invisibility of watermarks and robustness in the watermarking process. In the second stage, we employ generalized watermark representation learning to separate a unique representation of the watermark from the marked image in RGB space. Once trained, the HiWL model effectively learns generalizable watermark representations while maintaining broad applicability. Extensive experiments demonstrate the effectiveness of the proposed method. Specifically, it achieves 7.6% higher accuracy in watermark extraction compared to existing methods, while maintaining extremely low latency (processing 1000 images in 1 second).

图像水印深度学习版权保护高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。