提出无需训练的攻击方法,可高效移除多种扩散模型水印。
SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermark
- 通过随机重采样干扰生成轨迹,实现无感知水印移除
- 在9种水印方法上成功率95%~100%,视觉质量几乎无损
- 适用于噪声空间、频域和优化类水印,通用性强
基于扩散模型的水印方法通过操控初始噪声或反向扩散轨迹嵌入可验证标记。然而,这些方法均依赖于轨迹的精确重建以完成验证,这一假设构成根本性漏洞。本文提出训练无关的攻击方法SHIFT,利用随机扩散重采样在潜在空间中偏移生成轨迹,使重建图像在统计上与原始水印轨迹解耦,同时保持强视觉质量和语义一致性。在涵盖噪声空间、频域和优化类范式的九种代表性水印方法上,SHIFT实现95%至100%的攻击成功率,且几乎不损失语义质量,无需任何水印特异性知识或模型重训练。
原文摘要 · Abstract (English)
Diffusion-based watermarking methods embed verifiable marks by manipulating the initial noise or the reverse diffusion trajectory. However, these methods share a critical assumption: verification can succeed only if the diffusion trajectory can be faithfully reconstructed. This reliance on trajectory recovery constitutes a fundamental and exploitable vulnerability. We propose $\underline{\mathbf{S}}$tochastic $\underline{\mathbf{Hi}}$dden-Trajectory De$\underline{\mathbf{f}}$lec$\underline{\mathbf{t}}$ion ($\mathbf{SHIFT}$), a training-free attack that exploits this common weakness across diverse watermarking paradigms. SHIFT leverages stochastic diffusion resampling to deflect the generative trajectory in latent space, making the reconstructed image statistically decoupled from the original watermark-embedded trajectory while preserving strong visual quality and semantic consistency. Extensive experiments on nine representative watermarking methods spanning noise-space, frequency-domain, and optimization-based paradigms show that SHIFT achieves 95%--100% attack success rates with nearly no loss in semantic quality, without requiring any watermark-specific knowledge or model retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。