发现自回归图像生成水印易被移除和伪造,威胁内容溯源与数据清洗。
On the Robustness of Watermarking for Autoregressive Image Generation
- 提出三种新攻击:向量量化重生成、对抗优化、频率注入,均只需一张水印图
- 攻击者无需模型参数或密钥,即可成功移除水印并伪造虚假水印信号
- 现有水印方案不可靠,可能误判真实图像为合成内容,影响训练数据过滤
自回归图像生成器的普及迫切需要可靠的输出检测与归属机制,以遏制误导信息传播,并从训练数据中剔除合成图像,防止模型崩溃。为此,专门针对自回归模型设计的水印技术在生成时嵌入微弱信号,可通过对应检测器进行后续验证。本文研究此类方案,揭示其对水印移除与伪造攻击的脆弱性。我们评估了现有攻击方法,并提出三种新攻击:(i) 向量量化重生成移除攻击,(ii) 基于对抗优化的攻击,(iii) 频率注入攻击。实验表明,仅需一张水印参考图像且无需原始模型参数或水印密钥,即可有效实施移除与伪造攻击。结果表明,当前自回归图像生成水印方案无法可靠支持合成内容检测与数据集过滤。此外,它们还引发‘水印模仿’问题,即真实图像可被篡改以模仿生成器水印,导致误检并被排除在后续模型训练之外。
原文摘要 · Abstract (English)
The proliferation of autoregressive (AR) image generators demands reliable detection and attribution of their outputs to mitigate misinformation, and to filter synthetic images from training data to prevent model collapse. To address this need, watermarking techniques, specifically designed for AR models, embed a subtle signal at generation time, enabling downstream verification through a corresponding watermark detector. In this work, we study these schemes and demonstrate their vulnerability to both watermark removal and forgery attacks. We assess existing attacks and further introduce three new attacks: (i) a vector-quantized regeneration removal attack, (ii) adversarial optimization-based attack, and (iii) a frequency injection attack. Our evaluation reveals that removal and forgery attacks can be effective with access to a single watermarked reference image and without access to original model parameters or watermarking secrets. Our findings indicate that existing watermarking schemes for AR image generation do not reliably support synthetic content detection for dataset filtering. Moreover, they enable Watermark Mimicry, whereby authentic images can be manipulated to imitate a generator's watermark and trigger false detection to prevent their inclusion in future model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。