提出统一评估框架,量化判断图神经网络解释器的可信度。
Measuring What Matters: A Unified Evaluation Framework for GNN Explainability
- 将图结构与节点特征分开评估,避免依赖真实标签
- 发现无单一解释器在所有任务中表现最优
- 给出可落地的部署指南,助力实际应用
图可解释人工智能(G-XAI)对提升图神经网络的可解释性与可信度至关重要。尽管已有大量解释方法,但如何选择合适方法并评估其输出可靠性仍不明确。现有评估标准不一致,缺乏实用指导,限制了实际应用。本文提出一种无需真实标签假设的统一量化评估框架,将图数据的可解释性指标形式化为独立的拓扑结构与节点特征评估维度。大规模基准测试识别出在多个指标对和任务中始终位于帕累托前沿的解释器,确立了稳健的非支配解;同时证实不存在通用最优解释器。研究结果被提炼为可操作的G-XAI使用指南,帮助机器学习从业者评估并部署可信的GNN系统。
原文摘要 · Abstract (English)
Graph eXplainable AI (G-XAI) is increasingly important for making Graph Neural Networks interpretable and accountable. While a growing number of explainers are available, choosing the right method and assessing the trustworthiness of its outputs remains unclear. Consistent evaluation practices and actionable guidance are still missing, hindering practical adoption. In this paper, we introduce a unified, quantitative benchmarking framework for G-XAI that requires no ground-truth assumptions. We formalize tabular explainability metrics for graph data, evaluating topological structure and node features as independent components. Our large-scale benchmarking study identifies explainers that consistently lie on the Pareto front across metric pairs and tasks, establishing robustly non-dominated solutions - while confirming that no single explainer achieves universal superiority. We distill our findings into actionable G-XAI usability guidelines to support Machine Learning practitioners in evaluating and deploying trustworthy GNN-based pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。