arXiv:2505.16893stat.MLcs.LG2025-05被引 2

为图神经网络显著性图提供可靠统计检验,避免误判噪声为重要结构。

Statistical Test for Saliency Maps of Graph Neural Networks via Selective Inference

  • 基于选择性推断框架,解决数据重复使用导致的假阳性问题。
  • 在合成与真实数据上验证,能准确识别有意义的子图而非随机噪声。
  • 适用于多种分段线性显著性方法,提升解释结果可信度。

图神经网络(GNN)在处理各类图结构数据中表现突出,但其决策过程的可解释性仍是难题,因此常采用显著性图来识别关键节点与边组成的显著子图。然而,现有显著性图的可靠性受输入噪声影响,存在误判风险。本文提出一种基于选择性推断的统计检验框架,有效解决因数据双重使用导致的Ⅰ类错误率膨胀问题。该方法可生成统计有效的 $p$-值,在控制Ⅰ类错误率的同时,确保所识别的显著子图包含真实信息而非随机噪声。方法适用于具备分段线性特性的多种显著性方法(如类激活映射)。在合成数据与真实世界数据集上的实验表明,该框架能够可靠评估GNN解释结果的可信度。

原文摘要 · Abstract (English)

Graph Neural Networks (GNNs) have gained prominence for their ability to process graph-structured data across various domains. However, interpreting GNN decisions remains a significant challenge, leading to the adoption of saliency maps for identifying salient subgraphs composed of influential nodes and edges. Despite their utility, the reliability of GNN saliency maps has been questioned, particularly in terms of their robustness to input noise. In this study, we propose a statistical testing framework to rigorously evaluate the significance of saliency maps. Our main contribution lies in addressing the inflation of the Type I error rate caused by double-dipping of data, leveraging the framework of Selective Inference. Our method provides statistically valid $p$-values while controlling the Type I error rate, ensuring that identified salient subgraphs contain meaningful information rather than random artifacts. The method is applicable to a variety of saliency methods with piecewise linearity (e.g., Class Activation Mapping). We validate our method on synthetic and real-world datasets, demonstrating its capability in assessing the reliability of GNN interpretations.

图神经网络显著性图统计检验可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。