用图像风格特征做隐蔽后门攻击,让扩散模型生成恶意内容却难被发现。
Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models
- 用风格化特征作为隐蔽触发器,替代传统显眼图案或文字。
- 在图像修复任务中保持攻击效果,后门检测率低于0.5%。
- 适合研究模型安全、对抗攻击的人员阅读。
扩散模型(DMs)在图像生成中表现卓越,但近期研究揭示其易受后门攻击:攻击者通过在输入中嵌入隐蔽触发器操控输出。现有防御方法(如后门检测和触发器逆向)多因早期攻击依赖有限输入空间与低维、视觉明显的触发器而有效。为扩大威胁面,我们提出Gungnir,一种基于图像风格特征的新型后门攻击,利用风格化元素作为高阶、隐蔽的触发信号。引入重建对抗噪声(RAN)与短时步保留(STTR),确保图像到图像任务中触发器一致的扩散动态。生成样本在感知上与正常图像无异,可规避人工与自动检测。大量实验表明,Gungnir在不被检测的情况下绕过当前最先进防御,后门检测率(BDR)极低,并在微调净化后仍保持有效性,揭示了扩散模型中此前未被充分关注的安全漏洞。
原文摘要 · Abstract (English)
Diffusion Models (DMs) have achieved remarkable success in image generation, yet recent studies reveal their vulnerability to backdoor attacks, where adversaries manipulate outputs via covert triggers embedded in inputs. Existing defenses, such as backdoor detection and trigger inversion, are largely effective because prior attacks rely on limited input spaces and low-dimensional triggers that are visually conspicuous or easily captured by neural detectors. To broaden the threat landscape, we propose Gungnir, a novel backdoor attack that activates malicious behaviors through style-based triggers embedded in input images. Unlike explicit visual patches or textual cues, stylistic features serve as stealthy, high-level triggers. We introduce Reconstructing-Adversarial Noise (RAN) and Short-Term Timesteps-Retention (STTR) to preserve trigger-consistent diffusion dynamics in image-to-image tasks. The resulting trigger-embedded samples are perceptually indistinguishable from clean images, evading both manual and automated detection. Extensive experiments show that Gungnir bypasses state-of-the-art defenses with an extremely low backdoor detection rate (BDR) and remains effective under fine-tuning-based purification, revealing previously underexplored vulnerabilities in diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。