发现鲁棒水印反而易泄露信息,可被攻击者利用实现伪造和逃逸。
Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation
- 用预训练视觉模型提取多通道特征,挖掘水印冗余信息
- 仅需一张带水印图像,检测逃逸成功率提升60%
- 揭示鲁棒性与隐蔽性的矛盾,指导未来水印设计
水印在人工智能生成内容的溯源与检测中起关键作用。现有方法侧重于抵御真实世界失真(如JPEG压缩、加噪),但我们发现:此类鲁棒水印本质上会增加可检测模式的冗余度,造成可被利用的信息泄露。为此,我们提出一种攻击框架,通过预训练视觉模型的多通道特征学习,提取水印模式的泄漏信息。与以往需要大量数据或检测器访问的方法不同,本方法仅需一张水印图像即可实现伪造和检测逃逸。大量实验表明,相比当前最优方法,该方法在检测逃逸上提升60%成功率,在伪造精度上提高51%,同时保持视觉保真度。本工作揭示了鲁棒性与隐蔽性之间的悖论:当前“鲁棒”水印为抵抗失真牺牲了安全性,为未来水印设计提供了重要启示。
原文摘要 · Abstract (English)
Watermarking plays a key role in the provenance and detection of AI-generated content. While existing methods prioritize robustness against real-world distortions (e.g., JPEG compression and noise addition), we reveal a fundamental tradeoff: such robust watermarks inherently improve the redundancy of detectable patterns encoded into images, creating exploitable information leakage. To leverage this, we propose an attack framework that extracts leakage of watermark patterns through multi-channel feature learning using a pre-trained vision model. Unlike prior works requiring massive data or detector access, our method achieves both forgery and detection evasion with a single watermarked image. Extensive experiments demonstrate that our method achieves a 60\% success rate gain in detection evasion and 51\% improvement in forgery accuracy compared to state-of-the-art methods while maintaining visual fidelity. Our work exposes the robustness-stealthiness paradox: current "robust" watermarks sacrifice security for distortion resistance, providing insights for future watermark design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。