arXiv:2412.11441cs.CRcs.LG2024-12CVPR被引 16

用看不见的通用干扰让扩散模型偷偷执行恶意任务

UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models

  • 用通用对抗扰动生成无感触发器,适配任意图像和模型
  • 低中毒率下生成质量不降反升,攻击成功率超现有方法
  • 能骗过人眼和顶尖防御系统,适合研究安全风险

近期研究发现扩散模型易受后门攻击。现有攻击多使用明显可见的触发器(如灰色方块、眼镜),虽效果显著但易被人工或防御算法检测。若降低触发强度以增强隐蔽性,又会严重削弱攻击泛化性和有效性。本文提出UIBDiffusion,一种针对扩散模型的通用无感知后门攻击,可在保持优异攻击与生成性能的同时规避当前最先进防御。我们提出基于通用对抗扰动(UAPs)的新触发器生成方法,揭示此类初始为欺骗判别模型设计的扰动,可转化为扩散模型的强大无感知后门触发器。在多种扩散模型、采样器、数据集和目标上评估结果表明:1)通用性——单一触发器对任意图像和不同采样器的扩散模型均有效;2)实用性——在低中毒率下生成质量(如FID)相当甚至更优,攻击成功率(ASR)更高;3)不可检测性——触发器对人类感知自然,可绕过Elijah与TERD等主流扩散模型后门防御。代码与触发器将公开。

原文摘要 · Abstract (English)

Recent studies show that diffusion models (DMs) are vulnerable to backdoor attacks. Existing backdoor attacks impose unconcealed triggers (e.g., a gray box and eyeglasses) that contain evident patterns, rendering remarkable attack effects yet easy detection upon human inspection and defensive algorithms. While it is possible to improve stealthiness by reducing the strength of the backdoor, doing so can significantly compromise its generality and effectiveness. In this paper, we propose UIBDiffusion, the universal imperceptible backdoor attack for diffusion models, which allows us to achieve superior attack and generation performance while evading state-of-the-art defenses. We propose a novel trigger generation approach based on universal adversarial perturbations (UAPs) and reveal that such perturbations, which are initially devised for fooling pre-trained discriminative models, can be adapted as potent imperceptible backdoor triggers for DMs. We evaluate UIBDiffusion on multiple types of DMs with different kinds of samplers across various datasets and targets. Experimental results demonstrate that UIBDiffusion brings three advantages: 1) Universality, the imperceptible trigger is universal (i.e., image and model agnostic) where a single trigger is effective to any images and all diffusion models with different samplers; 2) Utility, it achieves comparable generation quality (e.g., FID) and even better attack success rate (i.e., ASR) at low poison rates compared to the prior works; and 3) Undetectability, UIBDiffusion is plausible to human perception and can bypass Elijah and TERD, the SOTA defenses against backdoors for DMs. We will release our backdoor triggers and code.

后门攻击扩散模型通用扰动安全风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。