用恶意图像攻击扩散模型,诱导生成不当内容
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion models
- 通过优化生成对抗图像,绕过文本过滤器
- 可成功突破Safe Latent Diffusion等防御机制
- 适用于研究模型安全与对抗攻击的学者
扩散模型在图像生成质量上取得显著进展,但也带来严重安全隐患,尤其容易生成不适宜工作场合(NSFW)内容。已有研究证明,通过对抗性文本提示可诱导生成此类内容,但这类提示常被基于文本的过滤器轻易检测,限制了其有效性。本文揭示了一种此前被忽视的漏洞:针对图像到图像(I2I)扩散模型的对抗性图像攻击。我们提出AdvI2I框架,通过优化生成器构造对抗图像,使扩散模型在不改变文本提示的情况下生成NSFW内容。该方法能有效绕过现有防御机制,如Safe Latent Diffusion(SLD)。此外,我们提出AdvI2I-Adaptive,该版本可适应潜在防御措施,并最小化对抗图像与NSFW概念嵌入之间的相似性,提升攻击鲁棒性。大量实验证明,两种方法均能有效突破当前防护体系,凸显加强I2I扩散模型安全防护的紧迫性。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly enhanced the quality of image synthesis, yet they have also introduced serious safety concerns, particularly the generation of Not Safe for Work (NSFW) content. Previous research has demonstrated that adversarial prompts can be used to generate NSFW content. However, such adversarial text prompts are often easily detectable by text-based filters, limiting their efficacy. In this paper, we expose a previously overlooked vulnerability: adversarial image attacks targeting Image-to-Image (I2I) diffusion models. We propose AdvI2I, a novel framework that manipulates input images to induce diffusion models to generate NSFW content. By optimizing a generator to craft adversarial images, AdvI2I circumvents existing defense mechanisms, such as Safe Latent Diffusion (SLD), without altering the text prompts. Furthermore, we introduce AdvI2I-Adaptive, an enhanced version that adapts to potential countermeasures and minimizes the resemblance between adversarial images and NSFW concept embeddings, making the attack more resilient against defenses. Through extensive experiments, we demonstrate that both AdvI2I and AdvI2I-Adaptive can effectively bypass current safeguards, highlighting the urgent need for stronger security measures to address the misuse of I2I diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。