提出可同时防御和放大扩散模型后门攻击的新框架
A Dual-Purpose Framework for Backdoor Defense and Backdoor Amplification in Diffusion Models
- 通过多步分布偏移与去噪一致性设计双重损失函数,精准反演触发器
- 检测准确率近乎完美,对复杂触发模式仍有效;攻击成功率接近100%
- 既可用于安全检测,也可帮助攻击者提升攻击效率
扩散模型虽在生成高质量多模态内容方面表现卓越,但近期研究揭示其易受后门攻击:当输入中嵌入预设触发器时,模型会生成特定有害输出(即后门目标)。本文提出PureDiffusion,一种兼具防御与攻击放大的双功能框架。防御方面,引入两种新损失函数:一是利用扩散过程多时间步的触发器诱导分布偏移,二是利用后门激活时的去噪一致性效应,实现触发器的高精度反演。基于反演结果,构建检测方法,分析反演触发器与生成的后门目标以识别攻击。在攻击角色下,该反演算法可强化原模型中的触发器,显著提升攻击性能,同时将训练时间减少最多20倍。实验表明,PureDiffusion检测准确率接近完美,远超现有防御方法,尤其针对复杂触发模式。攻击场景中,可使现有后门攻击成功率提升至近100%。
原文摘要 · Abstract (English)
Diffusion models have emerged as state-of-the-art generative frameworks, excelling in producing high-quality multi-modal samples. However, recent studies have revealed their vulnerability to backdoor attacks, where backdoored models generate specific, undesirable outputs called backdoor target (e.g., harmful images) when a pre-defined trigger is embedded to their inputs. In this paper, we propose PureDiffusion, a dual-purpose framework that simultaneously serves two contrasting roles: backdoor defense and backdoor attack amplification. For defense, we introduce two novel loss functions to invert backdoor triggers embedded in diffusion models. The first leverages trigger-induced distribution shifts across multiple timesteps of the diffusion process, while the second exploits the denoising consistency effect when a backdoor is activated. Once an accurate trigger inversion is achieved, we develop a backdoor detection method that analyzes both the inverted trigger and the generated backdoor targets to identify backdoor attacks. In terms of attack amplification with the role of an attacker, we describe how our trigger inversion algorithm can be used to reinforce the original trigger embedded in the backdoored diffusion model. This significantly boosts attack performance while reducing the required backdoor training time. Experimental results demonstrate that PureDiffusion achieves near-perfect detection accuracy, outperforming existing defenses by a large margin, particularly against complex trigger patterns. Additionally, in an attack scenario, our attack amplification approach elevates the attack success rate (ASR) of existing backdoor attacks to nearly 100\% while reducing training time by up to 20x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。