arXiv:2411.11688cs.CRcs.AI2024-11被引 3

给扩散模型加概念水印,能同时追踪图像和特定概念。

Watermarking Visual Concepts for Diffusion Models

  • 将目标概念与模型水印绑定,单次验证即可识别概念和溯源模型。
  • 在多个数据集上检测准确率提升6.3%至19.3%,优于现有方法。
  • 可抵御模型净化攻击,在恶意微调下仍保持21.7%的生成质量下降。

扩散模型的个性化生成能力虽强大,但易被用于生成未经授权的内容和虚假信息,威胁版权与网络安全。现有模型水印技术仅支持图像级溯源,缺乏概念可追溯性。当前方法需先检测概念再追溯模型,性能受限于概念识别准确率。本文提出轻量级概念水印框架ConceptWM,通过单阶段水印验证实现概念识别与模型溯源同步完成。为增强鲁棒性,设计一种协同嵌入水印的对抗扰动注入方法,防止模型净化攻击移除水印。实验表明,ConceptWM在COCO和StableDiffusionDB等多样数据集上检测准确率提升6.3%–19.3%。此外,其在对Stable Diffusion模型进行恶意微调(基于WikiArt和CelebA-HQ)时仍保持21.7%的FID/CLIP退化,有效抑制模型滥用。

原文摘要 · Abstract (English)

The personalization techniques of diffusion models succeed in generating images with specific concepts. This ability also poses great threats to copyright protection and network security since malicious users can generate unauthorized content and disinformation relevant to a target concept. Model watermarking is an effective solution to trace the malicious generated images and safeguard their copyright. However, existing model watermarking techniques merely achieve image-level tracing without concept traceability. When tracing infringing or harmful concepts, current approaches execute image concept detection and model tracing sequentially, where performance is critically constrained by concept detection accuracy. In this paper, we propose a lightweight concept watermarking framework that efficiently binds target concepts to model watermarks, supporting simultaneous concept identification and model tracing via single-stage watermark verification. To further enhance the robustness of concept watermarking, we propose an adversarial perturbation injection method collaboratively embedded with watermarks during image generation, avoiding watermark removal by model purification attacks. Experimental results demonstrate that ConceptWM significantly outperforms state-of-the-art watermarking methods, improving detection accuracy by 6.3%-19.3% across diverse datasets including COCO and StableDiffusionDB. Additionally, ConceptWM possesses a critical capability absent in other watermarking methods: it sustains a 21.7% FID/CLIP degradation under adversarial fine-tuning of Stable Diffusion models on WikiArt and CelebA-HQ, demonstrating its capability to mitigate model misuse.

水印扩散模型概念追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。