arXiv:2409.18897cs.CV2024-09被引 2

为文生图模型微调数据集设计水印,可追踪泄露并检测盗用。

Detecting Dataset Abuse in Fine-Tuning Stable Diffusion Models for Text-to-Image Synthesis

  • 在数据集微调时嵌入多策略水印,隐蔽性强。
  • 仅需修改2%数据即可实现高精度检测,不影响生成质量。
  • 水印具备鲁棒性和迁移性,适合版权保护场景。

文生图合成技术日益流行,常需使用领域特定数据集对生成模型进行微调以完成专业任务。然而,这些宝贵数据集面临未经授权使用和非法共享的风险,损害数据所有者的权益。本文针对稳定扩散模型在文生图微调过程中的数据集滥用问题,提出一种数据集水印框架,用于检测未经授权的使用并追溯数据泄露源头。该框架结合多种水印方案,具备大规模数据授权的有效性。大量实验表明,该方法仅需修改2%的数据即可实现高检测准确率,对数据集影响极小,且能有效追踪数据泄露路径。结果还验证了水印的鲁棒性与迁移能力,证明其在实际数据滥用检测中的可行性。

原文摘要 · Abstract (English)

Text-to-image synthesis has become highly popular for generating realistic and stylized images, often requiring fine-tuning generative models with domain-specific datasets for specialized tasks. However, these valuable datasets face risks of unauthorized usage and unapproved sharing, compromising the rights of the owners. In this paper, we address the issue of dataset abuse during the fine-tuning of Stable Diffusion models for text-to-image synthesis. We present a dataset watermarking framework designed to detect unauthorized usage and trace data leaks. The framework employs two key strategies across multiple watermarking schemes and is effective for large-scale dataset authorization. Extensive experiments demonstrate the framework's effectiveness, minimal impact on the dataset (only 2% of the data required to be modified for high detection accuracy), and ability to trace data leaks. Our results also highlight the robustness and transferability of the framework, proving its practical applicability in detecting dataset abuse.

数据水印文生图版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。