arXiv:2605.16796cs.CRcs.CV2026-05

用水印攻击水印,重加水印可高效消除原水印信号。

Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy

论文配图:Watermarks Attack Watermarks: Re-Watermarking as a Generic Removal Strategy
图 1 · 摘自论文原文
  • 通过在已加水印图像上再加水印,实现无需梯度的通用去水印。
  • 实验显示可使原始水印比特准确率下降25%至48%。
  • 可识别水印存在与类型,对现有防御构成严重威胁。

水印通过在输入图像中加入不可察觉的改动以触发检测器,用于溯源和保护知识产权。现有研究广泛关注对水印方案的攻击,因盗版或绕过深度伪造监管的动机强烈。本文观察到:攻击水印本质上也需对已加水印图像进行不可察觉的修改以触发检测器,这与水印机制高度相似。据此提出假设:水印可用于攻击水印。首个贡献在96种数据集、目标模型与攻击水印组合下验证该假设,结果表明简单重加水印即可可靠抑制原始信号,无需梯度、代理模型或检测密钥。第二个贡献是设计一个简单分类器,用于检测图像中是否存在及何种水印,实验显示准确率达0.878–0.953。结合两者,可使原始水印比特准确率降低至少25%,最高达48%。本方法成本低、通用性强,挑战了当前水印方案的可靠性,也质疑了复杂攻击的价值。

原文摘要 · Abstract (English)

Watermarking combines an imperceptible change to an input image that will trigger a detector, to assert provenance and protect intellectual property. The literature has shown great interest in attacks on watermarking schemes: attackers are clearly motivated to steal copyrighted material or circumvent legislated deepfake protections. In this work, we make a simple-yet-powerful observation: that such attacks on watermarking-like watermarks themselves-seek an imperceptible change to an input image (now already watermarked) that will trigger a detector. This analogy comparing watermark attacks to watermarking itself is highly suggestive: that watermarks could be used to attack watermarks. Our first contribution validates this hypothesis. In rigorous experiments spanning 96 combinations of dataset, victim, and attack watermarks, we show that simply re-watermarking an already watermarked image reliably suppresses the original signal, without requiring gradients, surrogate models, or detection keys. Our second contribution is a simple classifier for detecting the presence and identity of an existing watermark in a given image. Surprisingly, experimental findings demonstrate outstanding overall accuracies 0.878-0.953. This result is of independent interest as a security vulnerability: research shows that method-specific attacks achieve substantially stronger removal than black-box attacks. Taken together, watermark identification combined with re-watermarking successfully reduces bit accuracy by at least 25% and up to 48%. Our work constitutes a cheap, generic, and highly effective attack pipeline, calling into question the reliability of current watermarking schemes to such a simple attack, as well as the value of existing sophisticated attacks.

水印攻击通用攻击图像安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。