arXiv:2411.15367cs.CVcs.AI2024-11

用扩散模型反制图像水印保护,让训练数据被偷偷用。

Exploiting Watermark-Based Defense Mechanisms in Text-to-Image Diffusion Models for Unauthorized Data Usage

  • 利用扩散过程生成可控图像,避开水印细节但保留高层特征。
  • 仅需少量生成图像即可微调模型,成功率超90%。
  • 适合想绕过版权保护的研究者或攻击者参考。

文本到图像扩散模型(如Stable Diffusion)在生成高质量图像方面展现出巨大潜力。然而,近期研究指出,这些模型在训练中使用未经授权的数据可能引发知识产权侵犯或隐私问题。一种有前景的缓解方法是在图像上添加水印,并检查生成模型是否会复现类似水印特征。本文评估了多种基于水印的保护机制在文本到图像模型中的鲁棒性。我们发现,常见的图像变换对去除水印效果有限。因此,提出RATTAN方法:利用扩散过程对受保护输入进行可控图像生成,在保留输入高层特征的同时忽略水印所依赖的低层细节。随后,使用少量生成图像微调受保护模型。在三个数据集和140个文本到图像扩散模型上的实验表明,现有最先进保护措施无法抵御RATTAN攻击。

原文摘要 · Abstract (English)

Text-to-image diffusion models, such as Stable Diffusion, have shown exceptional potential in generating high-quality images. However, recent studies highlight concerns over the use of unauthorized data in training these models, which may lead to intellectual property infringement or privacy violations. A promising approach to mitigate these issues is to apply a watermark to images and subsequently check if generative models reproduce similar watermark features. In this paper, we examine the robustness of various watermark-based protection methods applied to text-to-image models. We observe that common image transformations are ineffective at removing the watermark effect. Therefore, we propose RATTAN, that leverages the diffusion process to conduct controlled image generation on the protected input, preserving the high-level features of the input while ignoring the low-level details utilized by watermarks. A small number of generated images are then used to fine-tune protected models. Our experiments on three datasets and 140 text-to-image diffusion models reveal that existing state-of-the-art protections are not robust against RATTAN.

图像生成水印攻击扩散模型数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。