arXiv:2508.11441cs.LGcs.AI2025-08被引 4

复杂模型的解释可能毫无意义,数学证明了这一点。

Informative Post-Hoc Explanations Only Exist for Simple Functions

  • 用学习理论定义解释有效性:能缩小可能函数空间
  • 梯度、SHAP等常见方法对复杂函数无效
  • 适合关注可解释性可信度的研究者与监管者

许多研究认为局部后处理解释算法可用于理解复杂机器学习模型的行为。然而,关于此类算法的理论保证仅存在于简单决策函数中,对于复杂模型是否成立尚不明确。本文提出一个基于学习理论的通用框架,定义解释需能降低可能决策函数的空间复杂度才算有效。基于此,我们证明:当应用于复杂函数时,多数流行解释算法实际上不具备信息量,从数学上否定了‘任何模型都可解释’的观点。我们进一步推导出不同算法变得有效的条件,通常比预期更强。例如,梯度解释和反事实解释在可微函数空间中无信息量;SHAP 和锚定解释在决策树空间中也不具备信息量。据此,我们讨论如何修改解释算法以使其有效。尽管分析为数学性质,但其对审计、监管及高风险AI应用具有深远实践意义。

原文摘要 · Abstract (English)

Many researchers have suggested that local post-hoc explanation algorithms can be used to gain insights into the behavior of complex machine learning models. However, theoretical guarantees about such algorithms only exist for simple decision functions, and it is unclear whether and under which assumptions similar results might exist for complex models. In this paper, we introduce a general, learning-theory-based framework for what it means for an explanation to provide information about a decision function. We call an explanation informative if it serves to reduce the complexity of the space of plausible decision functions. With this approach, we show that many popular explanation algorithms are not informative when applied to complex decision functions, providing a rigorous mathematical rejection of the idea that it should be possible to explain any model. We then derive conditions under which different explanation algorithms become informative. These are often stronger than what one might expect. For example, gradient explanations and counterfactual explanations are non-informative with respect to the space of differentiable functions, and SHAP and anchor explanations are not informative with respect to the space of decision trees. Based on these results, we discuss how explanation algorithms can be modified to become informative. While the proposed analysis of explanation algorithms is mathematical, we argue that it holds strong implications for the practical applicability of these algorithms, particularly for auditing, regulation, and high-risk applications of AI.

可解释性理论分析模型审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。