arXiv:2506.06018cs.MMcs.AI2025-06被引 4

无需优化即可通用伪造扩散模型水印,实现高保真攻击

Optimization-Free Universal Watermark Forgery with Regenerative Diffusion Models

  • 利用再生扩散模型直接提取并植入水印,不需额外优化
  • 24种组合下水印可检测率最高达100%,视觉质量保持优秀
  • 适合关注生成内容安全与水印防御的研究者和从业者

水印是追踪和验证人工智能生成图像来源的关键技术,但存在风险。现有研究已证明可在不了解目标生成模型和水印方案的情况下,通过对抗优化伪造水印。本文揭示一种更严重的无优化、通用性水印伪造风险:利用现有的再生扩散模型,提出PnP(Plug-and-Plant)攻击方法,通过图像再生过程无缝提取并集成目标水印,无需额外优化。该方法独立于目标图像来源或水印模型,以目标图像的水印潜在表示和覆盖图像的视觉-文本上下文作为先验引导再生采样。在24种模型-数据-水印组合下的广泛评估表明,PnP能成功伪造水印(最高可检测率达100%且用户可追溯),同时保持最佳视觉感知。该方法绕过模型重训练,具备对任意图像的适应性,显著扩大了伪造攻击范围,对当前扩散模型水印技术的安全性及水印在合成数据生成与治理中的权威性构成更大挑战。

原文摘要 · Abstract (English)

Watermarking becomes one of the pivotal solutions to trace and verify the origin of synthetic images generated by artificial intelligence models, but it is not free of risks. Recent studies demonstrate the capability to forge watermarks from a target image onto cover images via adversarial optimization without knowledge of the target generative model and watermark schemes. In this paper, we uncover a greater risk of an optimization-free and universal watermark forgery that harnesses existing regenerative diffusion models. Our proposed forgery attack, PnP (Plug-and-Plant), seamlessly extracts and integrates the target watermark via regenerating the image, without needing any additional optimization routine. It allows for universal watermark forgery that works independently of the target image's origin or the watermarking model used. We explore the watermarked latent extracted from the target image and visual-textual context of cover images as priors to guide sampling of the regenerative process. Extensive evaluation on 24 scenarios of model-data-watermark combinations demonstrates that PnP can successfully forge the watermark (up to 100% detectability and user attribution), and maintain the best visual perception. By bypassing model retraining and enabling adaptability to any image, our approach significantly broadens the scope of forgery attacks, presenting a greater challenge to the security of current watermarking techniques for diffusion models and the authority of watermarking schemes in synthetic data generation and governance.

水印伪造扩散模型生成安全再生模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。