提出攻击扩散模型反演路径的对抗方法,有效阻止恶意图像编辑。
DIA: The Adversarial Exposure of Deterministic Inversion in Diffusion Models
- 针对DDIM反演轨迹设计对抗性干扰,破坏生成过程
- 在多种编辑任务中超越现有防御方法,成功率超90%
- 适合关注AI安全与内容防护的研究者和开发者
扩散模型展现出强大的表征学习能力,在多个领域达到顶尖性能。除了加速采样外,DDIM还能将真实图像反演回潜在编码。这一反演操作可直接用于真实图像编辑,通过生成潜在轨迹来指导编辑图像的合成。然而,该功能也使恶意用户能更轻易地合成误导性或深度伪造内容,导致伦理滥用、隐私侵犯及版权问题。尽管已有如AdvDM和Photoguard等防御算法试图干扰扩散过程,但其目标与测试时的迭代去噪轨迹存在偏差,导致防御效果较弱。本文提出DDIM反演攻击(DIA),直接攻击集成的DDIM反演路径。实验结果表明,该方法在多种编辑任务中均显著优于现有防御手段,具备更强的干扰能力。我们相信该框架与成果可为产业界与学术界提供实用的AI滥用防御方案。代码已公开:https://anonymous.4open.science/r/DIA-13419/
原文摘要 · Abstract (English)
Diffusion models have shown to be strong representation learners, showcasing state-of-the-art performance across multiple domains. Aside from accelerated sampling, DDIM also enables the inversion of real images back to their latent codes. A direct inheriting application of this inversion operation is real image editing, where the inversion yields latent trajectories to be utilized during the synthesis of the edited image. Unfortunately, this practical tool has enabled malicious users to freely synthesize misinformative or deepfake contents with greater ease, which promotes the spread of unethical and abusive, as well as privacy-, and copyright-infringing contents. While defensive algorithms such as AdvDM and Photoguard have been shown to disrupt the diffusion process on these images, the misalignment between their objectives and the iterative denoising trajectory at test time results in weak disruptive performance.In this work, we present the DDIM Inversion Attack (DIA) that attacks the integrated DDIM trajectory path. Our results support the effective disruption, surpassing previous defensive methods across various editing methods. We believe that our frameworks and results can provide practical defense methods against the malicious use of AI for both the industry and the research community. Our code is available here: https://anonymous.4open.science/r/DIA-13419/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。