arXiv:2605.16815cs.CRcs.LG2026-05

提出新方法防御图神经网络中的隐蔽后门攻击。

Universal Graph Backdoor Defense: A Feature-based Homophily Perspective

论文配图:Universal Graph Backdoor Defense: A Feature-based Homophily Perspective
图 1 · 摘自论文原文
  • 从特征同质性角度分析攻击共性,发现后门节点局部特征不一致。
  • 通过邻居感知重建损失检测并清除后门,攻击成功率下降90%以上。
  • 对新型特征类后门有效,适合高风险场景的模型安全加固。

图神经网络在关系学习中表现卓越,但易受图后门攻击(GBA)威胁,限制其在高风险场景的应用。现有防御方法多针对基于子图的攻击,依赖中毒目标节点与触发子图显式连接的假设。我们实证发现,此类结构中心方法无法应对保留图拓扑的特征类后门攻击。本文提出通用图后门防御框架,从特征同质性视角揭示:无论触发机制如何,后门节点的局部特征一致性均显著低于正常节点。据此,我们设计基于邻居感知重构损失的节点级特征一致性度量,用于识别后门节点,并结合鲁棒训练策略消除触发影响、降低检测不确定性带来的噪声。大量实验表明,该方法在子图与特征类攻击下均显著降低攻击成功率(最高达90%+),同时保持良好的干净精度。

原文摘要 · Abstract (English)

Graph neural networks (GNNs) have achieved remarkable success in relational learning. However, their vulnerability to graph backdoor attacks (GBAs) poses a significant barrier to broader adoption in high-stakes applications. Despite recent advances in graph backdoor defense (GBD), existing methods primarily focus on subgraph-based GBAs, relying on the assumption that poisoned target nodes are explicitly connected to subgraph triggers. Our empirical results reveal that such structure-centric approaches fail to defend against emerging feature-based GBAs that preserve graph topology. Therefore, in this paper, we study a novel problem of universal graph backdoor defense. First, we investigate the shared effects of both attack types from a feature-based homophily perspective, which characterizes local feature consistency between nodes and their neighborhoods. Thorough theoretical and empirical analyses demonstrate that, regardless of trigger mechanisms, backdoors induced by GBAs exhibit lower feature-based homophily than clean nodes, indicating a discrepancy in local feature similarity. Motivated by this insight, we propose to leverage node-level local feature consistency, modeled by a neighbor-aware reconstruction loss, to distinguish backdoors from clean nodes. Then, a robust training strategy is developed to eliminate trigger effects while reducing noise induced by detection uncertainty. Extensive experiments demonstrate that our framework significantly degrades the attack success rate and maintains competitive clean accuracy under both subgraph-based and feature-based attacks.

图神经网络后门防御特征同质性安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。