arXiv:2508.07636cs.LGcs.AI2025-08TPAMI被引 5

从理论视角梳理解释深度模型的方法,解决其可信度难题。

Attribution Explanations for Deep Neural Networks: A Theoretical Perspective

  • 提出三大理论挑战:方法杂乱、缺乏根基、难验证真伪
  • 构建统一框架,揭示不同方法的共性与差异
  • 为选择和设计更可信的解释方法提供理论支持

归因解释是解释深度神经网络(DNN)的典型方法,通过为每个输入变量分配重要性或贡献分数来理解输出决策。近年来涌现出大量归因方法,但核心问题仍未解决:哪些方法真正反映输入变量对决策的实际贡献?该可信性问题严重影响解释的可靠性与实用性。本文指出三大根源挑战:第一,方法间结构差异大,启发式、形式化与实现方式不统一;第二,多数方法缺乏坚实的理论基础,其合理性模糊或未被验证;第三,缺乏真实标签,难以实证评估可信度。近期理论进展为此提供了新路径,本文系统总结三方面关键方向:(i) 理论统一,揭示方法间的共性与差异,实现系统比较;(ii) 理论溯源,阐明现有方法的逻辑基础;(iii) 理论评估,严格证明方法是否满足可信性原则。除全面综述外,本文还探讨这些研究如何深化理论认知、指导方法选择并启发新方法设计,并展望未来有前景的开放问题。

原文摘要 · Abstract (English)

Attribution explanation is a typical approach for explaining deep neural networks (DNNs), inferring an importance or contribution score for each input variable to the final output. In recent years, numerous attribution methods have been developed to explain DNNs. However, a persistent concern remains unresolved, i.e., whether and which attribution methods faithfully reflect the actual contribution of input variables to the decision-making process. The faithfulness issue undermines the reliability and practical utility of attribution explanations. We argue that these concerns stem from three core challenges. First, difficulties arise in comparing attribution methods due to their unstructured heterogeneity, differences in heuristics, formulations, and implementations that lack a unified organization. Second, most methods lack solid theoretical underpinnings, with their rationales remaining absent, ambiguous, or unverified. Third, empirically evaluating faithfulness is challenging without ground truth. Recent theoretical advances provide a promising way to tackle these challenges, attracting increasing attention. We summarize these developments, with emphasis on three key directions: (i) Theoretical unification, which uncovers commonalities and differences among methods, enabling systematic comparisons; (ii) Theoretical rationale, clarifying the foundations of existing methods; (iii) Theoretical evaluation, rigorously proving whether methods satisfy faithfulness principles. Beyond a comprehensive review, we provide insights into how these studies help deepen theoretical understanding, inform method selection, and inspire new attribution methods. We conclude with a discussion of promising open problems for further work.

归因解释深度学习理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。