用逆向风格迁移生成图数据反事实解释,更准确且结构更合理。
Graph Inverse Style Transfer for Counterfactual Explainability
- 将反事实生成看作逆向过程,利用谱风格迁移保持结构和语义一致
- 在8个基准上提升反事实有效性7.6%,对真实类别解释力提升45.5%
- 适合研究图神经网络可解释性的研究人员,尤其关注决策边界分析
反事实解释旨在通过最小化输入变化来揭示模型决策原因。对于图数据,该任务因需保持结构完整性和语义一致性而尤为困难。不同于以往依赖前向扰动的方法,本文提出图逆向风格迁移(GIST),首次将图反事实生成重构为回溯过程,利用谱风格迁移技术。通过匹配原始输入的全局结构谱特征并保持局部内容忠实性,GIST生成的反事实是输入风格与反事实内容之间的插值。在8个二分类和多分类图分类基准上,GIST实现反事实有效率提升7.6%,对真实类别分布的解释力显著增强45.5%。此外,其回溯机制有效避免越过预测器决策边界,大幅减小输入与反事实间的谱差异。这些结果挑战了传统前向扰动方法,为图可解释性提供了新视角。
原文摘要 · Abstract (English)
Counterfactual explainability seeks to uncover model decisions by identifying minimal changes to the input that alter the predicted outcome. This task becomes particularly challenging for graph data due to preserving structural integrity and semantic meaning. Unlike prior approaches that rely on forward perturbation mechanisms, we introduce Graph Inverse Style Transfer (GIST), the first framework to re-imagine graph counterfactual generation as a backtracking process, leveraging spectral style transfer. By aligning the global structure with the original input spectrum and preserving local content faithfulness, GIST produces valid counterfactuals as interpolations between the input style and counterfactual content. Tested on 8 binary and multi-class graph classification benchmarks, GIST achieves a remarkable +7.6% improvement in the validity of produced counterfactuals and significant gains (+45.5%) in faithfully explaining the true class distribution. Additionally, GIST's backtracking mechanism effectively mitigates overshooting the underlying predictor's decision boundary, minimizing the spectral differences between the input and the counterfactuals. These results challenge traditional forward perturbation methods, offering a novel perspective that advances graph explainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。