arXiv:2602.21593cs.LGcs.CR2026-02中稿 · The Web Conference…被引 5

用大模型精准改语义,让内容水印失效

Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection

  • 用大模型引导语义修改,保持图像整体一致
  • 在不破坏视觉连贯性的前提下成功擦除水印
  • 揭示当前语义水印对大模型攻击的脆弱性

生成图像在社交媒体和网络版权分发场景中泛滥,语义水印被越来越多地集成到扩散模型中,以支持可靠的来源追踪和伪造防范。传统基于噪声层的水印仍易受反向攻击,可恢复嵌入信号。为缓解此问题,近期的内容感知语义水印将水印信号绑定于高层图像语义,限制会破坏全局一致性的局部编辑。然而,大语言模型(LLM)具备结构化推理能力,可实现针对语义空间的精准探索,实现局部细微但全局一致的语义改动,从而破坏此类绑定。为揭示这一被忽视的漏洞,我们提出一种保持连贯性的语义注入(CSI)攻击,利用大模型引导的语义操作,并在嵌入空间相似性约束下进行。该对齐机制在保持视觉-语义一致性的同时,选择性扰动与水印相关的语义,最终导致检测器误判。大量实验证明,CSI持续优于现有攻击基线,在对抗内容感知语义水印方面表现更优,揭示了当前语义水印设计在面对大模型驱动的语义扰动时存在根本性安全弱点。

原文摘要 · Abstract (English)

Generative images have proliferated on Web platforms in social media and online copyright distribution scenarios, and semantic watermarking has increasingly been integrated into diffusion models to support reliable provenance tracking and forgery prevention for web content. Traditional noise-layer-based watermarking, however, remains vulnerable to inversion attacks that can recover embedded signals. To mitigate this, recent content-aware semantic watermarking schemes bind watermark signals to high-level image semantics, constraining local edits that would otherwise disrupt global coherence. Yet, large language models (LLMs) possess structured reasoning capabilities that enable targeted exploration of semantic spaces, allowing locally fine-grained but globally coherent semantic alterations that invalidate such bindings. To expose this overlooked vulnerability, we introduce a Coherence-Preserving Semantic Injection (CSI) attack that leverages LLM-guided semantic manipulation under embedding-space similarity constraints. This alignment enforces visual-semantic consistency while selectively perturbing watermark-relevant semantics, ultimately inducing detector misclassification. Extensive empirical results show that CSI consistently outperforms prevailing attack baselines against content-aware semantic watermarking, revealing a fundamental security weakness of current semantic watermark designs when confronted with LLM-driven semantic perturbations.

语义水印大模型攻击图像安全扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。