提出一种无需依赖模型的水印去除方法,可无损移除各类AI图像水印。
MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation

- 通过潜空间统计关联性识别并解耦水印信号
- 平均比特准确率降至53.14%,接近随机猜测
- 适用于多种水印系统,且对画质影响极小
数字水印已成为AI生成图像溯源与版权归属的关键技术,但其在真实场景下抵御模型无关攻击的鲁棒性仍缺乏研究。现有攻击要么仅针对特定生成模型,要么导致严重视觉失真。本文提出MarkNull,一种基于潜空间流形操纵的模型无关水印移除方法。核心发现:带水印图像的生成潜表示与嵌入的初始噪声存在强统计依赖性。为此,我们引入噪声-潜表示对齐得分(NLAS),构建优化目标,选择性解耦潜表示与水印信号,同时保持语义保真度。在后处理、微调和初始噪声三类水印方案上广泛评估表明,MarkNull将平均比特准确率降至53.14%,接近随机猜测(50%),且无可见质量下降。为进一步提升效率,提出MarkNull-A——一种无需优化的单前向传播版本,实现0.50秒/图的处理速度,计算开销轻微。值得注意的是,该攻击成功破解了Google SynthID-Image系统,并可有效迁移至视频水印。最后,我们设计了一种攻击检测机制作为防御对照,强调需开发能抵御模型无关潜空间攻击的水印方案。
原文摘要 · Abstract (English)
Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustness against realistic, model-agnostic removal attacks remains poorly explored. Existing attacks either succeed only against specific generative models or achieve removal at the cost of severe visual degradation. In this paper, we propose MarkNull, a model-agnostic watermark removal attack via on-manifold latent manipulation. MarkNull is grounded in a key observation: watermarked images exhibit a strong statistical dependency between the generated latent representation and the embedded initial noise. To quantify this dependency, we introduce the Noise-Latent Alignment Score (NLAS) and formulate an optimization objective that selectively decorrelates the latent representation from the embedded watermark while preserving semantic fidelity. Extensive evaluations across different categories of watermarking paradigms, including post-hoc, fine-tuning-based, and initial-noise-based schemes, demonstrate that MarkNull reduces average bit accuracy to 53.14%, approaching random-guessing (50%), without perceptible image degradation. To further improve scalability, we propose MarkNull-A, an amortized, optimization-free variant that distills the attack into a single forward pass, achieving 0.50 s/image with modest computational overhead. Notably, our attacks successfully compromise Google's SynthID-Image system while preserving high visual quality and transfer effectively to video watermarking. Finally, we present an attack detection mechanism as a defensive counterpart to MarkNull and MarkNull-A, highlighting the necessity of developing watermark designs resilient to model-agnostic latent-space attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。