arXiv:2508.21072cs.CV2025-08被引 9

提出两种攻破隐形水印的方法,实测移除率超95%且图像质量几乎无损。

First-Place Solution to NeurIPS 2024 Invisible Watermark Removal Challenge

  • 针对不同攻击场景,分别采用自适应VAE和扩散模型进行水印剥离。
  • 在黑白盒测试中实现95.7%的水印移除率,图像保真度高。
  • 适合研究水印安全性的学者及数字版权保护开发者参考。

内容水印是数字媒体认证与版权保护的重要工具,但其对抗性攻击下的鲁棒性尚不明确。我们提交了NeurIPS 2024「擦除隐形水印」挑战赛的冠军解决方案,该挑战在不同攻击者知识条件下测试水印鲁棒性。挑战分为黑盒与灰盒两个赛道:对于灰盒场景,我们采用基于VAE的自适应逃逸攻击,结合测试时优化与CIELAB空间的颜色对比度恢复,以保持图像质量;对于黑盒场景,我们首先根据空间或频域中的图像伪影对图像进行聚类,随后对每类使用带可控噪声注入与来自ChatGPT生成描述的语义先验的图像到图像扩散模型,并优化参数设置。实证评估表明,该方法在近乎完美地移除水印(95.7%)的同时,对残留图像质量影响极小。我们希望这些攻击能推动更鲁棒图像水印技术的发展。

原文摘要 · Abstract (English)

Content watermarking is an important tool for the authentication and copyright protection of digital media. However, it is unclear whether existing watermarks are robust against adversarial attacks. We present the winning solution to the NeurIPS 2024 Erasing the Invisible challenge, which stress-tests watermark robustness under varying degrees of adversary knowledge. The challenge consisted of two tracks: a black-box and beige-box track, depending on whether the adversary knows which watermarking method was used by the provider. For the beige-box track, we leverage an adaptive VAE-based evasion attack, with a test-time optimization and color-contrast restoration in CIELAB space to preserve the image's quality. For the black-box track, we first cluster images based on their artifacts in the spatial or frequency-domain. Then, we apply image-to-image diffusion models with controlled noise injection and semantic priors from ChatGPT-generated captions to each cluster with optimized parameter settings. Empirical evaluations demonstrate that our method successfully achieves near-perfect watermark removal (95.7%) with negligible impact on the residual image's quality. We hope that our attacks inspire the development of more robust image watermarking methods.

水印攻击扩散模型图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。