用反向触发机制检测并反制生成模型中的后门攻击
PureDiffusion: Using Backdoor to Counter Backdoor in Generative Diffusion Models
- 通过逆向生成后门触发器来检测扩散模型中的隐蔽攻击
- 在多种触发-目标组合下,逆向触发器保真度和攻击成功率均显著优于现有方法
- 特别适用于需要高精度防御的生成式AI系统安全评估
扩散模型(DMs)在众多生成任务中表现出色,但近期研究揭示其易受后门攻击:当输入包含特定触发器时,模型会持续生成预设的恶意结果(如有害图像)。尽管已有多种攻击手段被提出,针对此类威胁的防御方法仍有限且研究不足,尤其是在逆向还原触发器方面。本文提出PureDiffusion,一种新型后门防御框架,能高效通过逆向生成嵌入在扩散模型中的后门触发器来检测攻击。大量实验表明,在多种触发器-目标对上,PureDiffusion在保真度(即逆向触发器与原始触发器的相似度)和后门成功率(即逆向触发器引发对应后门目标的比率)方面均显著优于现有方法。值得注意的是,在某些情况下,PureDiffusion逆向生成的触发器甚至比原始触发器具有更高的攻击成功率。
原文摘要 · Abstract (English)
Diffusion models (DMs) are advanced deep learning models that achieved state-of-the-art capability on a wide range of generative tasks. However, recent studies have shown their vulnerability regarding backdoor attacks, in which backdoored DMs consistently generate a designated result (e.g., a harmful image) called backdoor target when the models' input contains a backdoor trigger. Although various backdoor techniques have been investigated to attack DMs, defense methods against these threats are still limited and underexplored, especially in inverting the backdoor trigger. In this paper, we introduce PureDiffusion, a novel backdoor defense framework that can efficiently detect backdoor attacks by inverting backdoor triggers embedded in DMs. Our extensive experiments on various trigger-target pairs show that PureDiffusion outperforms existing defense methods with a large gap in terms of fidelity (i.e., how much the inverted trigger resembles the original trigger) and backdoor success rate (i.e., the rate that the inverted trigger leads to the corresponding backdoor target). Notably, in certain cases, backdoor triggers inverted by PureDiffusion even achieve higher attack success rate than the original triggers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。