arXiv:2509.10577cs.CRcs.AI2025-09中稿 · IEEE EuroS&P 2026被引 5

揭示生成模型水印在对抗篡改下的理论极限,证明超过一半比特被修改就无法可靠检测。

The Coding Limits of Robust Watermarking for Generative Models

  • 提出零比特篡改检测编码模型,形式化鲁棒水印的核心需求
  • 理论证明:二进制情况下,超过50%比特被篡改则水印必失效
  • 实验验证:裁剪缩放操作可精准翻转约一半隐空间符号,抹除水印

我们研究生成模型密码水印的一个基本问题:当攻击者可篡改编码信号时,水印能保持多可靠?为此,我们引入一种最小编码抽象——零比特篡改检测码。这是一种密钥机制,用于采样伪随机码字,并判断候选码字是否为未标记内容或已被篡改的有效码字。该模型捕捉了鲁棒水印的两个核心要求:正确性与篡改检测。在该抽象下,我们证明了对独立符号篡改的严格无条件极限:对于大小为$q$的字母表,存在一个临界篡改率$1-1/q$,一旦攻击者修改的符号比例超过此值,即使允许对随机内容有固定常数级误报率,也无法可靠检测篡改。特别地,在二进制情况下,任何密码水印在超过一半编码比特被修改后均无法保持鲁棒性。我们还通过简单的信息论构造证明该阈值是紧的,可在所有严格小于该率的篡改情况下实现正确性与篡改检测。随后,我们通过实验检验该极限是否在现实中出现,针对Gunn、Zhao和Song(ICLR 2025)提出的图像水印方法进行测试,结果表明,简单的裁剪与缩放操作能可靠地翻转约一半的潜在符号,并持续阻止信念传播解码恢复码字,从而在视觉上保持图像不变的同时彻底擦除水印。

原文摘要 · Abstract (English)

We study a basic question about cryptographic watermarking for generative models: how reliable can a watermark remain when an adversary is allowed to corrupt the encoded signal? To address this question, we introduce a minimal coding abstraction that we call a zero-bit tamper-detection code. This is a secret-key procedure that samples a pseudorandom codeword and, given a candidate word, decides whether it should be treated as unmarked content or as the result of tampering with a valid codeword. It captures the two core requirements of robust watermarking: soundness and tamper detection. Within this abstraction we prove a sharp unconditional limit on robustness to independent symbol corruption. For an alphabet of size $q$, there is a critical corruption rate of $1-1/q$ such that no scheme with soundness, even relaxed to allow a fixed constant false positive probability on random content, can reliably detect tampering once an adversary can change more than this fraction of symbols. In particular, in the binary case no cryptographic watermark can remain robust if more than half of the encoded bits are modified. We also show that this threshold is tight by giving simple information-theoretic constructions that achieve soundness and tamper detection for all strictly smaller corruption rates. We then test experimentally whether this limit appears in practice by looking at the recent watermarking for images of Gunn, Zhao, and Song (ICLR 2025). We show that a simple crop and resize operation reliably flipped about half of the latent signs and consistently prevented belief-propagation decoding from recovering the codeword, erasing the watermark while leaving the image visually intact.

水印安全生成模型信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。