arXiv:2511.09332cs.LGcs.AI2025-11中稿 · AAAI被引 74

提出基于数据分布的特征归因方法,让黑箱模型决策更可解释。

Distribution-Based Feature Attribution for Explaining the Predictions of Any Classifier

  • 从数据分布出发定义特征归因问题,确保解释有统计基础。
  • 实验表明新方法比现有最优方法更准确高效。
  • 适合需要可信解释的AI应用,如医疗、金融决策系统。

复杂黑箱AI模型的普及加剧了对其决策解释的需求。特征归因方法已成为提供事后解释的流行方案,但该领域长期缺乏正式的问题定义。本文通过引入特征归因的形式化定义来填补这一空白,要求解释必须基于给定数据集所代表的潜在概率分布。分析发现,许多现有模型无关方法不满足此条件,即使满足也常存在其他局限。为此,我们提出分布型特征归因解释(DFAX),一种新型的模型无关特征归因方法。DFAX是首个直接基于数据分布解释分类器预测的方法。通过大量实验验证,其在有效性与效率上均优于当前最先进基线方法。

原文摘要 · Abstract (English)

The proliferation of complex, black-box AI models has intensified the need for techniques that can explain their decisions. Feature attribution methods have become a popular solution for providing post-hoc explanations, yet the field has historically lacked a formal problem definition. This paper addresses this gap by introducing a formal definition for the problem of feature attribution, which stipulates that explanations be supported by an underlying probability distribution represented by the given dataset. Our analysis reveals that many existing model-agnostic methods fail to meet this criterion, while even those that do often possess other limitations. To overcome these challenges, we propose Distributional Feature Attribution eXplanations (DFAX), a novel, model-agnostic method for feature attribution. DFAX is the first feature attribution method to explain classifier predictions directly based on the data distribution. We show through extensive experiments that DFAX is more effective and efficient than state-of-the-art baselines.

可解释AI特征归因黑箱模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。