从基础模型出发构建特征归因框架,无需强加理论约束。
Feature Attribution from First Principles
- 从最简单的指示函数出发,逐步构建复杂模型的归因方法
- 推导出深度ReLU网络的归因闭式表达式
- 为优化归因评估指标提供新思路,适合模型可解释性研究者
特征归因方法广泛用于解释机器学习模型的行为,通过为每个输入特征分配重要性分数来量化其对模型预测的影响。然而,这些方法的实证评估仍面临重大挑战。为克服这一缺陷,已有研究提出基于公理的框架,要求所有归因方法必须满足特定条件。本文认为此类公理往往过于严格,因此提出一种从零开始构建的新归因框架:不预先设定公理,而是从最简单的指示函数定义归因,并将其作为构建复杂模型归因的基础。我们证明,根据原子归因的不同选择,可恢复多种现有归因方法。进一步地,我们推导出深度ReLU网络的归因闭式表达式,并朝着基于归因优化评估指标的方向迈出一步。
原文摘要 · Abstract (English)
Feature attribution methods are a popular approach to explain the behavior of machine learning models. They assign importance scores to each input feature, quantifying their influence on the model's prediction. However, evaluating these methods empirically remains a significant challenge. To bypass this shortcoming, several prior works have proposed axiomatic frameworks that any feature attribution method should satisfy. In this work, we argue that such axioms are often too restrictive, and propose in response a new feature attribution framework, built from the ground up. Rather than imposing axioms, we start by defining attributions for the simplest possible models, i.e., indicator functions, and use these as building blocks for more complex models. We then show that one recovers several existing attribution methods, depending on the choice of atomic attribution. Subsequently, we derive closed-form expressions for attribution of deep ReLU networks, and take a step toward the optimization of evaluation metrics with respect to feature attributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。