统一解释方法框架,让模型归因更可比、可验证。
Unifying Attribution-Based Explanations Using Functional Decomposition
- 用特征移除思想统一现有归因方法
- 证明所有可加分解都源于同一构造原理
- 为解释方法提供理论保障,适合研究者与工程师
机器学习的黑箱问题催生了大量解释方法,但方法差异大导致选择困难。本文提出统一框架,首次引入基于移除的归因方法(RBAMs),表明众多现有方法均可视为此类。提出规范加性分解(CAD),基于移除特征的思想构建任意函数的可加分解。证明所有有效可加分解均为CAD实例,且每种移除型方法对应特定CAD。进一步揭示每种移除型方法本质是特定合作博弈下的博弈论值或交互指数,该博弈由对应CAD定义。利用此内在联系,提出形式化行为描述(功能公理),并给出满足这些公理的充分条件。最后展示如何基于该框架高效近似已有解释方法。
原文摘要 · Abstract (English)
The black box problem in machine learning has led to the introduction of an ever-increasing set of explanation methods for complex models. These explanations have different properties, which in turn has led to the problem of method selection: which explanation method is most suitable for a given use case? In this work, we propose a unifying framework of attribution-based explanation methods, which provides a step towards a rigorous study of the similarities and differences of explanations. We first introduce removal-based attribution methods (RBAMs), and show that an extensively broad selection of existing methods can be viewed as such RBAMs. We then introduce the canonical additive decomposition (CAD). This is a general construction for additively decomposing any function based on the central idea of removing (groups of) features. We proceed to show that indeed every valid additive decomposition is an instance of the CAD, and that any removal-based attribution method is associated with a specific CAD. Next, we show that any removal-based attribution method can be completely defined as a game-theoretic value or interaction index for a specific (possibly constant-shifted) cooperative game, which is defined using the corresponding CAD of the method. We then use this intrinsic connection to define formal descriptions of specific behaviours of explanation methods, which we also call functional axioms, and identify sufficient conditions on the corresponding CAD and game-theoretic value or interaction index of an attribution method under which the attribution method is guaranteed to adhere to these functional axioms. Finally, we show how this unifying framework can be used to develop new, efficient approximations for existing explanation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。