arXiv:2505.00808cs.LGcs.AI2025-05AAAI被引 8

为神经网络的机制可解释性提供数学哲学基础,强调因果解释的可提取性与可信度。

A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i

  • 提出解释忠实性概念,量化解释与模型间的匹配程度。
  • 定义机制可解释性为可验证的、本体论级的因果机制解释。
  • 主张解释乐观原则是机制可解释性成功的必要前提,适用于理论研究者。

机制可解释性旨在通过因果解释理解神经网络。我们支持解释观点假说:机制可解释性研究是一种有原则的方法,因为神经网络中包含可提取和理解的隐式解释。因此,我们证明了解释忠实性(Explanatory Faithfulness)这一评估指标是明确定义的。我们提出将机制可解释性(MI)定义为生成模型层面、本体论、因果机制且可证伪的解释实践,从而将其与其他可解释性范式区分开,并揭示其内在局限性。我们还提出了解释乐观原则,认为这是机制可解释性成功所必需的前提假设。

原文摘要 · Abstract (English)

Mechanistic Interpretability aims to understand neural networks through causal explanations. We argue for the Explanatory View Hypothesis: that Mechanistic Interpretability research is a principled approach to understanding models because neural networks contain implicit explanations which can be extracted and understood. We hence show that Explanatory Faithfulness, an assessment of how well an explanation fits a model, is well-defined. We propose a definition of Mechanistic Interpretability (MI) as the practice of producing Model-level, Ontic, Causal-Mechanistic, and Falsifiable explanations of neural networks, allowing us to distinguish MI from other interpretability paradigms and detail MI's inherent limits. We formulate the Principle of Explanatory Optimism, a conjecture which we argue is a necessary precondition for the success of Mechanistic Interpretability.

机制可解释性因果解释神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。