为微调扩散模型设计水印评估框架,发现现有方法易被移除。
Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach
- 构建统一威胁模型与三维度评估框架
- 多数水印在真实攻击下仍可被完全移除
- 适合关注生成模型版权追踪的研究者
近期扩散模型的微调技术可复现特定图像(如人脸或艺术风格),但也带来版权与安全风险。数据集水印通过在训练图像中嵌入不可察觉的标记,使输出仍可追溯,但现有方法缺乏统一评估体系。本文提出通用威胁模型,并建立涵盖普适性、传递性与鲁棒性的综合评估框架。实验表明,现有方法在普适性与传递性上表现良好,对常见图像处理具备一定鲁棒性,但在真实威胁场景下仍显脆弱。为此,论文进一步提出一种实用的水印移除方法,可在不干扰微调的前提下完全消除水印,揭示了未来研究的关键挑战。
原文摘要 · Abstract (English)
Recent fine-tuning techniques for diffusion models enable them to reproduce specific image sets, such as particular faces or artistic styles, but also introduce copyright and security risks. Dataset watermarking has been proposed to ensure traceability by embedding imperceptible watermarks into training images, which remain detectable in outputs even after fine-tuning. However, current methods lack a unified evaluation framework. To address this, this paper establishes a general threat model and introduces a comprehensive evaluation framework encompassing Universality, Transmissibility, and Robustness. Experiments show that existing methods perform well in universality and transmissibility, and exhibit some robustness against common image processing operations, yet still fall short under real-world threat scenarios. To reveal these vulnerabilities, the paper further proposes a practical watermark removal method that fully eliminates dataset watermarks without affecting fine-tuning, highlighting a key challenge for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。