用语义引导的局部重绘,可无痕擦除主流隐形水印
Removing Watermarks with Partial Regeneration using Semantic Information
- 基于视觉语言模型生成细粒度描述,用零样本分割提取前景,仅重绘背景
- 在4种水印系统上均成功擦除水印,保留图像感知质量(掩码SSIM达0.94)
- 适合研究水印安全性或对抗攻击的从业者,揭示现有防御的语义漏洞
随着AI生成图像日益普及,隐形水印成为版权与来源验证的主要手段。最新水印方案嵌入语义信号——内容感知的模式,旨在抵御常见图像操作,但其对自适应攻击者的鲁棒性尚未充分探索。本文揭示了一种此前未被报告的漏洞,并提出SemanticRegen,一种三阶段、无需标签的攻击方法,在保持图像语义不变的前提下,可有效消除最先进的语义型与隐形水印。该流程首先利用视觉语言模型获取细粒度描述,其次通过零样本分割提取前景掩码,最后使用大语言模型引导的扩散模型仅对背景进行修复。在1000个提示下评估四种水印系统(TreeRing、StegaStamp、StableSig、DWT/DCT),SemanticRegen是唯一能在统计上击败语义型TreeRing水印的方法(p=0.10>0.05),其余方案比特准确率均低于0.75,同时保持高感知质量(掩码SSIM=0.94±0.01)。我们还引入掩码SSIM(mSSIM)量化前景区域保真度,结果显示本方法比以往基于扩散模型的攻击高出最多12%。这些结果凸显当前水印防御与自适应语义感知攻击能力之间的巨大差距,亟需开发能抵御内容保持型再生攻击的新水印算法。
原文摘要 · Abstract (English)
As AI-generated imagery becomes ubiquitous, invisible watermarks have emerged as a primary line of defense for copyright and provenance. The newest watermarking schemes embed semantic signals - content-aware patterns that are designed to survive common image manipulations - yet their true robustness against adaptive adversaries remains under-explored. We expose a previously unreported vulnerability and introduce SemanticRegen, a three-stage, label-free attack that erases state-of-the-art semantic and invisible watermarks while leaving an image's apparent meaning intact. Our pipeline (i) uses a vision-language model to obtain fine-grained captions, (ii) extracts foreground masks with zero-shot segmentation, and (iii) inpaints only the background via an LLM-guided diffusion model, thereby preserving salient objects and style cues. Evaluated on 1,000 prompts across four watermarking systems - TreeRing, StegaStamp, StableSig, and DWT/DCT - SemanticRegen is the only method to defeat the semantic TreeRing watermark (p = 0.10 > 0.05) and reduces bit-accuracy below 0.75 for the remaining schemes, all while maintaining high perceptual quality (masked SSIM = 0.94 +/- 0.01). We further introduce masked SSIM (mSSIM) to quantify fidelity within foreground regions, showing that our attack achieves up to 12 percent higher mSSIM than prior diffusion-based attackers. These results highlight an urgent gap between current watermark defenses and the capabilities of adaptive, semantics-aware adversaries, underscoring the need for watermarking algorithms that are resilient to content-preserving regenerative attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。