提出频域调制扩散攻击框架,有效消除图像水印且保持画质。
Breaking Watermarks in the Frequency Domain: A Modulated Diffusion Attack Framework

- 在扩散过程的正向与反向阶段引入频域水印调制模块。
- 攻击后图像保真度优于现有方法,跨水印方案泛化能力强。
- 适合研究水印防御与攻击的学者,尤其关注生成模型安全者。
数字图像水印技术在生成式AI版权保护中迅速发展,但水印攻击技术进展相对有限,打破了攻防平衡并阻碍了该领域进一步发展。本文提出FMDiffWA,一种基于频域调制的扩散攻击框架。具体而言,引入频域水印调制(FWM)模块,并将其嵌入扩散过程的正向与反向采样阶段。该机制可选择性调节与水印相关的频域成分,从而有效消除隐形水印信号,同时保持被攻击图像的感知质量。为在攻击效果与视觉保真度间取得更好权衡,我们通过在标准噪声估计目标基础上增加辅助精炼约束,重构了传统扩散模型的训练策略。大量实验表明,FMDiffWA在视觉保真度上优于现有水印攻击方法,且对多种水印方案具有强泛化能力。
原文摘要 · Abstract (English)
Digital image watermarking has advanced rapidly for copyright protection of generative AI, yet the comparatively limited progress in watermark attack techniques has broken the attack-defense balance and hindered further advances in the field. In this paper, we propose FMDiffWA, a frequency-domain modulated diffusion framework for watermark attacks. Specifically, we introduce a frequency-domain watermark modulation (FWM) module and incorporate it into the sampling stages both the forward and reverse diffusion processes. This mechanism enables selective modulation of watermark-related frequency components, thereby allowing FMDiffWA to effectively neutralize the invisible watermark signals while preserving the perceptual quality of the attacked watermarked images. To achieve a better trade-off between attack efficacy and visual fidelity, we reformulate the training strategy of conventional diffusion models by augmenting the canonical noise estimation objective with an auxiliary refinement constraint. Comprehensive experiments demonstrate that FMDiffWA achieves superior visual fidelity compared to existing watermark attacks, while exhibiting strong generalization across diverse watermarking schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。