HMARK在扩散模型中实现可迁移的多比特语义水印,保护训练数据版权。
HMARK: Radioactive Multi-Bit Semantic-Latent Watermarking for Diffusion Models
- 将版权信息编码至图像扩散模型的语义潜空间(h-space)
- 水印检测准确率达98.57%,比特恢复率95.07%,无漏检
- 适合用于版权保护的扩散模型,尤其对微调后输出有效
现代生成式扩散模型依赖大规模训练数据集,其中常包含权属或使用权限不明的图像。放射性水印——即能传递至模型输出的标记——有助于检测未经授权的数据是否被用于训练。此外,有效的水印还需具备不可感知性、鲁棒性和多比特容量等特性。为此,我们提出HMARK,一种新型多比特水印方案,将所有权信息以秘密比特形式嵌入图像扩散模型的语义潜空间(h-space)。通过利用h-space的可解释性与语义意义,确保水印信号对应有意义的语义属性,使得HMARK嵌入的水印具有放射性、对各类失真具备鲁棒性,且对视觉质量影响极小。实验结果表明,HMARK在下游对抗性模型经LoRA微调后生成的图像上,实现了98.57%的水印检测准确率、95.07%的比特级恢复准确率、100%的召回率和1.0的AUC值。
原文摘要 · Abstract (English)
Modern generative diffusion models rely on vast training datasets, often including images with uncertain ownership or usage rights. Radioactive watermarks -- marks that transfer to a model's outputs -- can help detect when such unauthorized data has been used for training. Moreover, aside from being radioactive, an effective watermark for protecting images from unauthorized training also needs to meet other existing requirements, such as imperceptibility, robustness, and multi-bit capacity. To overcome these challenges, we propose HMARK, a novel multi-bit watermarking scheme, which encodes ownership information as secret bits in the semantic-latent space (h-space) for image diffusion models. By leveraging the interpretability and semantic significance of h-space, ensuring that watermark signals correspond to meaningful semantic attributes, the watermarks embedded by HMARK exhibit radioactivity, robustness to distortions, and minimal impact on perceptual quality. Experimental results demonstrate that HMARK achieves 98.57% watermark detection accuracy, 95.07% bit-level recovery accuracy, 100% recall rate, and 1.0 AUC on images produced by the downstream adversarial model finetuned with LoRA on watermarked data across various types of distortions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。