arXiv:2503.06340cs.CRcs.LG2025-03被引 3

首次揭示离散图扩散模型的后门攻击风险,威胁药物设计等关键应用。

Backdoor Attacks on Discrete Graph Diffusion Models

  • 设计新型后门攻击,可隐蔽植入恶意生成模式。
  • 攻击后模型仍能生成高质量图,且后门触发持久有效。
  • 适用于关注图生成安全性的研究人员与开发者。

扩散模型在图像、视频等连续数据领域表现强大,而离散图扩散模型(DGDMs)最近被用于图生成,已在分子、蛋白质建模等领域达到当前最优性能。然而,在缺乏对安全漏洞认知的情况下,将此类模型部署于药物发现等安全敏感场景存在风险。本文首次系统研究了针对图扩散模型的后门攻击——一种同时影响训练与推理/生成阶段的严重威胁。我们定义了攻击威胁模型,设计出的攻击能使受控模型在未激活后门时生成高质量图,激活后则生成有效、隐蔽且持久的恶意图。此外,所生成图满足置换不变性与可交换性,这两项是图生成模型的核心特性。第1、2点通过有无防御的实证评估验证,第3点通过理论分析确认。

原文摘要 · Abstract (English)

Diffusion models are powerful generative models in continuous data domains such as image and video data. Discrete graph diffusion models (DGDMs) have recently extended them for graph generation, which are crucial in fields like molecule and protein modeling, and obtained the SOTA performance. However, it is risky to deploy DGDMs for safety-critical applications (e.g., drug discovery) without understanding their security vulnerabilities. In this work, we perform the first study on graph diffusion models against backdoor attacks, a severe attack that manipulates both the training and inference/generation phases in graph diffusion models. We first define the threat model, under which we design the attack such that the backdoored graph diffusion model can generate 1) high-quality graphs without backdoor activation, 2) effective, stealthy, and persistent backdoored graphs with backdoor activation, and 3) graphs that are permutation invariant and exchangeable--two core properties in graph generative models. 1) and 2) are validated via empirical evaluations without and with backdoor defenses, while 3) is validated via theoretical results.

图生成安全攻击扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。