给图像加隐形干扰噪声,防被恶意生成伪造内容
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
- 用统一损失函数对抗多种参考生成技术
- 可防御灰盒模型和部分商用API,攻击效果强
- 适合关注图像安全与隐私保护的研究者
扩散模型在生成高质量图像方面表现卓越,但其滥用可能导致虚假新闻或针对个人的有害内容,引发严重社会问题。本文提出Anti-Reference,一种通过向图像添加不可察觉的对抗性噪声来抵御基于参考生成技术威胁的新方法。我们设计了一种统一损失函数,可同时攻击基于微调的定制化方法、非微调定制化方法以及以人类为中心的驱动方法。基于该损失,训练对抗噪声编码器预测噪声,或直接使用PGD方法优化噪声。实验表明,该方法具备一定的迁移攻击能力,能有效应对灰盒模型及部分商业API。大量实验证明了Anti-Reference的有效性,确立了图像安全领域的新基准。
原文摘要 · Abstract (English)
Diffusion models have revolutionized generative modeling with their exceptional ability to produce high-fidelity images. However, misuse of such potent tools can lead to the creation of fake news or disturbing content targeting individuals, resulting in significant social harm. In this paper, we introduce Anti-Reference, a novel method that protects images from the threats posed by reference-based generation techniques by adding imperceptible adversarial noise to the images. We propose a unified loss function that enables joint attacks on fine-tuning-based customization methods, non-fine-tuning customization methods, and human-centric driving methods. Based on this loss, we train a Adversarial Noise Encoder to predict the noise or directly optimize the noise using the PGD method. Our method shows certain transfer attack capabilities, effectively challenging both gray-box models and some commercial APIs. Extensive experiments validate the performance of Anti-Reference, establishing a new benchmark in image security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。