arXiv:2606.29807cs.CRcs.CV2026-06中稿 · ICML

从几何畸变视角揭示语义水印伪造的内在瓶颈

Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective

论文配图:Rethinking Forgery Attacks on Semantic Watermarks in Black-Box Settings: A Geometric Distortion Perspective
图 1 · 摘自论文原文
  • 将水印伪造问题建模为潜在空间中的率失真问题
  • 发现代理与目标模型结构不匹配导致不可逾越的失真下限
  • 提出无需依赖具体方案的伪造检测方法,适用于多种黑盒场景

近期研究显示,嵌入在隐空间扩散模型初始噪声中的语义水印在黑盒环境下易受伪造攻击。然而,现有方法多依赖经验证据,缺乏对攻击成败条件的严谨理论理解。为此,我们从潜在空间的率失真角度重新审视此类攻击。分析表明,由于代理模型与目标模型间的结构不匹配,存在不可消除的失真下限,从根本上限制了伪造水印的保真度。我们进一步将该失真表征为潜在流形上的结构性几何偏移,表现为全局漂移与局部形变,而非随机噪声。基于此,我们提出一种方案无关的检测方法,可在水印验证前识别伪造样本。大量实验验证了该方法在多种黑盒场景下的有效性,同时保持对常见失真的鲁棒性。

原文摘要 · Abstract (English)

Recent studies have shown that semantic watermarks, which embed information into the initial noise of latent diffusion models (LDMs), are vulnerable to black-box forgery attacks. However, existing methods primarily rely on empirical evidence and lack a rigorous theoretical understanding of the conditions under which such attacks succeed or fail. To bridge this gap, we rethink the nature of such attacks through the lens of rate-distortion in the latent space. Our analysis identifies an irreducible distortion floor due to structural mismatches between proxy and target models, which fundamentally limits the fidelity of forged watermarks. We further characterize this distortion as structured geometric deviations on the latent manifold, in the form of global drift and local deformation rather than stochastic noise. Leveraging these insights, we propose a scheme-agnostic detection method that distinguishes forged samples before watermark verification. Extensive experiments demonstrate the effectiveness of our method across diverse black-box scenarios, while preserving robustness to common distortions.

水印安全扩散模型黑盒攻击几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。