arXiv:2507.10643stat.MLcs.AI2025-07

提出基于泰勒展开的可调节归因方法,提升黑箱模型解释的可信度与实用性。

TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models

  • 基于泰勒展开构建归因框架,确保特征贡献分配符合逻辑原则。
  • 在多个数据集和模型上验证,归因结果更贴合用户目标,且解释清晰。
  • 适用于可微与不可微模型,适合需要高可信解释的AI应用者。

后处理模型无关局部归因(LA)方法广泛用于解释黑箱AI模型,通过量化特征贡献来提升可解释性。然而,现有方法多依赖启发式或部分合理的归因机制,且归因质量常受下游任务目标影响,缺乏统一标准。本文提出泰勒展开起源的自适应归因方法(TaylorPODA),基于泰勒展开框架,建立一组公理化要求,明确特征与泰勒项间的精确对应关系。分析表明,现有方法在理论严谨性与用户目标适配间存在根本矛盾。TaylorPODA引入可控的交互效应分配机制,使归因结果能适应用户定义的目标,同时满足所有公理。此外,该方法还具备哈桑尼红利解释,可拓展至非可微模型。理论证明其满足所有公理及适应性属性。实验证明,在多个数据集和可微/不可微模型上,TaylorPODA均显著提升归因结果与用户目标的一致性,且保持良好可读性。本工作为构建更可信的可解释AI系统提供了基础。

原文摘要 · Abstract (English)

Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise contributions. However, many existing methods rely on heuristic or only partially justified attribution mechanisms, while the quality of attribution itself is often shaped by downstream objectives without universally accepted standards. In this work, we propose Taylor exPansion-Originated aDaptive Attribution (TaylorPODA), a new post-hoc model-agnostic LA method grounded in the Taylor expansion framework. We first introduce a set of postulates, which formalize principled requirements for explicitly and exhaustively attributing Taylor terms to the corresponding features. Based on these postulates, we analyze existing post-hoc model-agnostic LA methods and identify a fundamental tension between principled attribution and adaptation toward user-defined utilities. To address this challenge, TaylorPODA introduces a controllable allocation mechanism for Taylor interaction effects, enabling attribution results to adapt to downstream objectives while preserving the proposed postulates. Furthermore, although developed from a Taylor-expansion perspective, TaylorPODA also admits a Harsanyi-dividend interpretation, allowing the attribution mechanism to extend beyond model differentiability. Theoretical analysis demonstrates that TaylorPODA satisfies all the proposed postulates together with an additional adaptation property. Empirical results across multiple datasets and both differentiable and non-differentiable models further show that TaylorPODA achieves consistently improved alignment with user-defined utilities while maintaining the communicability of the resulting explanations. Overall, this work provides a starting point toward more trustworthy XAI systems for the deployment of increasingly powerful yet opaque task models.

可解释AI归因方法泰勒展开

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。