提出对抗性攻击破坏生成图像水印,提升防御能力。
RoboSignature: Robust Signature and Watermarking on Network Attacks
- 用对抗微调攻击破坏扩散模型的水印嵌入能力。
- 新防御方法使水印在攻击下仍可检测,鲁棒性显著提升。
- 适合关注生成内容安全与防伪造的研究者参考。
生成模型使得仅凭一个提示即可轻松创建各类图像,但也引发了关于内容真实来源的伦理担忧。为区分人工创作与模型生成内容,对生成数据进行水印成为主流方法,目标是让所有生成图像隐含不可见水印,便于后续检测或识别。当前的稳定水印(Stable Signature)通过微调潜在扩散模型(LDM)解码器,使生成图像中嵌入唯一水印。本文提出一种新型对抗性微调攻击,能有效干扰模型嵌入预期水印的能力,揭示现有水印方法的重大漏洞。为应对该问题,我们进一步设计了一种受扰抗性微调算法,借鉴大语言模型中的方法,针对LDM水印需求进行优化。研究结果强调了在生成系统中预见并防御潜在漏洞的重要性。
原文摘要 · Abstract (English)
Generative models have enabled easy creation and generation of images of all kinds given a single prompt. However, this has also raised ethical concerns about what is an actual piece of content created by humans or cameras compared to model-generated content like images or videos. Watermarking data generated by modern generative models is a popular method to provide information on the source of the content. The goal is for all generated images to conceal an invisible watermark, allowing for future detection or identification. The Stable Signature finetunes the decoder of Latent Diffusion Models such that a unique watermark is rooted in any image produced by the decoder. In this paper, we present a novel adversarial fine-tuning attack that disrupts the model's ability to embed the intended watermark, exposing a significant vulnerability in existing watermarking methods. To address this, we further propose a tamper-resistant fine-tuning algorithm inspired by methods developed for large language models, tailored to the specific requirements of watermarking in LDMs. Our findings emphasize the importance of anticipating and defending against potential vulnerabilities in generative systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。