arXiv:2511.05598cs.CReess.IV2025-11被引 4

扩散模型编辑会破坏鲁棒隐写水印,导致信息几乎完全丢失。

Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks

  • 利用扩散模型的迭代去噪过程削弱水印信号
  • 对多种水印方案攻击后解码准确率接近零
  • 适合关注生成式AI安全风险的研究者

鲁棒隐形水印旨在将隐藏信息嵌入图像中,使其在各类操作后仍可检测且不可见。然而,强大的基于扩散的图像生成与编辑模型能实现内容保留的逼真变换,却可能无意中擦除或扭曲嵌入的水印。本文通过理论与实证分析表明,扩散模型编辑可有效破坏当前主流的鲁棒水印技术。我们分析了扩散模型迭代加噪与去噪过程如何衰减水印信号,并提供形式化证明:在特定条件下,生成图像几乎无法检测到水印信息。基于此,我们提出一种扩散驱动攻击,通过生成重制来擦除图像中的水印;进一步引入引导扩散攻击,将水印解码器嵌入采样循环,直接针对水印进行破坏。我们在多个近期深度学习水印方案(如 StegaStamp、TrustMark、VINE)上评估,结果表明扩散编辑可将水印解码准确率降至近零,同时保持图像高视觉保真度。研究揭示了当前鲁棒水印技术在生成式模型编辑下的根本脆弱性,凸显了在生成式AI时代亟需新型水印策略。

原文摘要 · Abstract (English)

Robust invisible watermarking aims to embed hidden messages into images such that they survive various manipulations while remaining imperceptible. However, powerful diffusion-based image generation and editing models now enable realistic content-preserving transformations that can inadvertently remove or distort embedded watermarks. In this paper, we present a theoretical and empirical analysis demonstrating that diffusion-based image editing can effectively break state-of-the-art robust watermarks designed to withstand conventional distortions. We analyze how the iterative noising and denoising process of diffusion models degrades embedded watermark signals, and provide formal proofs that under certain conditions a diffusion model's regenerated image retains virtually no detectable watermark information. Building on this insight, we propose a diffusion-driven attack that uses generative image regeneration to erase watermarks from a given image. Furthermore, we introduce an enhanced \emph{guided diffusion} attack that explicitly targets the watermark during generation by integrating the watermark decoder into the sampling loop. We evaluate our approaches on multiple recent deep learning watermarking schemes (e.g., StegaStamp, TrustMark, and VINE) and demonstrate that diffusion-based editing can reduce watermark decoding accuracy to near-zero levels while preserving high visual fidelity of the images. Our findings reveal a fundamental vulnerability in current robust watermarking techniques against generative model-based edits, underscoring the need for new watermarking strategies in the era of generative AI.

隐写水印扩散模型生成安全攻击方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。