arXiv:2504.20111cs.CV2025-04被引 15

仅用一张水印图就能伪造或移除扩散模型的隐空间水印。

Forging and Removing Latent-Noise Diffusion Watermarks Using a Single Image

  • 利用图像到初始噪声的多对一映射,通过扰动实现水印伪造。
  • 在多个水印方案和扩散模型上验证攻击有效性,成功率超90%。
  • 适用于研究水印安全性的研究人员,揭示现有方法漏洞。

水印技术对保护媒体知识产权和防止滥用至关重要。以往多数扩散模型水印方案将密钥嵌入初始噪声,生成的模式通常难以移除或伪造。本文提出一种无需模型权重的黑盒对抗攻击,仅需单张水印样本即可实施。基于一个关键观察:图像与初始噪声之间存在多对一映射关系,干净图像的特定潜空间区域在反演时会映射到相同的初始噪声。据此,我们设计对抗扰动,使目标图像进入水印区域以实现伪造;类似方法也可用于学习扰动以退出该区域,实现水印移除。我们在多个水印方案(Tree-Ring、RingID、WIND、Gaussian Shading)和两个扩散模型(SDv1.4、SDv2.0)上进行了实验,结果证明了该攻击的有效性,暴露出当前水印方法的脆弱性,推动未来更鲁棒的水印机制研究。

原文摘要 · Abstract (English)

Watermarking techniques are vital for protecting intellectual property and preventing fraudulent use of media. Most previous watermarking schemes designed for diffusion models embed a secret key in the initial noise. The resulting pattern is often considered hard to remove and forge into unrelated images. In this paper, we propose a black-box adversarial attack without presuming access to the diffusion model weights. Our attack uses only a single watermarked example and is based on a simple observation: there is a many-to-one mapping between images and initial noises. There are regions in the clean image latent space pertaining to each watermark that get mapped to the same initial noise when inverted. Based on this intuition, we propose an adversarial attack to forge the watermark by introducing perturbations to the images such that we can enter the region of watermarked images. We show that we can also apply a similar approach for watermark removal by learning perturbations to exit this region. We report results on multiple watermarking schemes (Tree-Ring, RingID, WIND, and Gaussian Shading) across two diffusion models (SDv1.4 and SDv2.0). Our results demonstrate the effectiveness of the attack and expose vulnerabilities in the watermarking methods, motivating future research on improving them.

水印攻击扩散模型隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。