arXiv:2601.21610cs.CV2026-01

用视觉语言模型统一评估扩散模型图像水印,兼顾可解释性与安全性。

WMVLM: Evaluating Diffusion Model Image Watermarking via Vision-Language Models

  • 基于视觉语言模型构建统一评估框架,区分残留与语义水印。
  • 在多个数据集和模型上超越现有方法,通用性强。
  • 支持可解释的文本生成,帮助理解水印效果与攻击行为。

数字水印对保护扩散模型生成图像至关重要。准确的水印评估对于算法开发不可或缺,但现有方法存在诸多局限:缺乏残留与语义水印的统一评估框架,结果不可解释,忽视综合安全考量,且对语义水印使用不当指标。为此,我们提出 WMVLM——首个基于视觉语言模型(VLMs)的统一且可解释的扩散模型图像水印评估框架。针对不同水印类型重新定义质量与安全度量:残留水印通过伪影强度与擦除抗性评估,语义水印则通过潜在分布偏移衡量。我们设计三阶段训练策略,逐步实现分类、评分与可解释文本生成。实验表明,WMVLM 在跨数据集、扩散模型与水印方法上均优于当前最优 VLM,具备强泛化能力。

原文摘要 · Abstract (English)

Digital watermarking is essential for securing generated images from diffusion models. Accurate watermark evaluation is critical for algorithm development, yet existing methods have significant limitations: they lack a unified framework for both residual and semantic watermarks, provide results without interpretability, neglect comprehensive security considerations, and often use inappropriate metrics for semantic watermarks. To address these gaps, we propose WMVLM, the first unified and interpretable evaluation framework for diffusion model image watermarking via vision-language models (VLMs). We redefine quality and security metrics for each watermark type: residual watermarks are evaluated by artifact strength and erasure resistance, while semantic watermarks are assessed through latent distribution shifts. Moreover, we introduce a three-stage training strategy to progressively enable the model to achieve classification, scoring, and interpretable text generation. Experiments show WMVLM outperforms state-of-the-art VLMs with strong generalization across datasets, diffusion models, and watermarking methods.

图像水印视觉语言模型扩散模型可解释评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。