提出PGID框架,提升扩散模型水印检测抗攻击能力。
PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection

- 通过渐进式反演与去噪,将被篡改的隐空间投影回原始区域。
- 在多个攻击场景下恢复丢失水印,识别伪造图像,准确率超90%。
- 无需训练、可即插即用,适合版权保护与内容安全应用。
随着AI生成图像的普及,数字水印成为保护知识产权和防范恶意滥用的重要手段。现有语义水印方法虽能高效保护扩散模型版权,但依赖扩散反演进行水印检测,存在关键漏洞。伪造攻击通过将带水印隐变量移至无水印区域,同时引导无水印隐变量进入水印区域,实现欺骗。我们分析发现,此类攻击成功源于隐空间的偏移。为此,提出首个无需训练、可即插即用的噪声提取框架PGID,通过渐进式反演-去噪循环,消除中间隐变量扰动,缓解对抗性干扰,将被扰动的隐变量精确投影回其原始所属区域。跨多种攻击方案的全面评估表明,PGID有效恢复水印检测可靠性,成功复原被移除水印并识别伪造实例。
原文摘要 · Abstract (English)
With the proliferation of AI-generated images, digital watermarking has become an essential safeguard for protecting intellectual property and mitigating malicious exploitation. Recent works on semantic watermarking have enabled efficient copyright protection for diffusion models. However, the dependence of semantic watermarking on diffusion inversion for watermark detection creates a critical vulnerability. Imprint removal and forgery attacks exploit this weakness to produce deceptive results. Our analysis reveals that these attacks succeed by displacing watermarked latents into the unwatermarked region, while guiding unwatermarked latents into the watermarked region. Based on that, we propose Progressive Guided Inversion and Denoising (PGID), the first plug-and-play, training-free noise extraction framework designed to defend against both attack strategies. PGID effectively defends by projecting perturbed latents back to the region where they originally belong. The projection is achieved by eliminating intermediate latent deflections and mitigating adversarial perturbations through progressive inversion-denoising cycles. Comprehensive evaluations across multiple schemes demonstrate that PGID successfully restores detection reliability by recovering removed watermarks and identifying forged instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。