arXiv:2607.14271cs.LGcs.AI2026-07被引 1

统一解释AI的特征归因方法,揭示其数学假设与常见错误。

Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist

  • 从五个关键选择构建统一框架,涵盖主流归因方法
  • 揭示基线敏感、对抗攻击等失败源于特定数学假设
  • 提出10项报告清单,提升归因研究透明度

特征归因方法是可解释人工智能的核心。其假设以多种数学形式表达:合作博弈值、路径积分、梯度算子、扰动分布和反向传播规则。本文提出一个局部可加特征归因的统一框架,围绕五项规范选择——值函数、参考点、路径、扰动分布与守恒规则——整合了Shapley值、路径法、梯度/反向传播、扰动法及CAM类方法。通过方法-公理矩阵对比各方法,并将基线敏感性、离流形扰动、合理性检验失败、对抗操纵及方法分歧等常见缺陷,追溯至其对应的假设根源。最后,提出一项包含十项内容的报告检查清单,强调归因结果仅在特定数学假设下有意义,且这些假设必须明确报告。

原文摘要 · Abstract (English)

Feature-attribution methods are central to explainable artificial intelligence. Their assumptions are expressed in several mathematical languages: cooperative-game values, path integrals, gradient operators, perturbation distributions, and backpropagation rules. This survey proposes a common framework for local additive feature attribution. It organizes Shapley, path-based, gradient/backpropagation, perturbation, and CAM-style methods around five specification choices: value function, reference, path, perturbation distribution, and conservation rule. It then compares these methods through an axiom-by-method matrix and links common failure modes, including baseline sensitivity, off-manifold perturbations, sanity-check failures, adversarial manipulation, and method disagreement, to the assumptions that produce them. Finally, the survey proposes a ten-item reporting checklist for studies that use local additive attributions. The central message is that attribution results are meaningful only relative to the mathematical assumptions under which they are defined, and that those assumptions should be reported.

可解释AI特征归因数学框架报告清单

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。