提出四维哲学框架,评估神经网络解释的优劣
Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
- 从科学哲学四大视角构建解释评价体系
- 紧凑证明法综合多种优点,潜力巨大
- 适合研究可解释AI与模型机制的学者
机制可解释性(MI)旨在通过因果解释理解神经网络。尽管已有多种生成解释的方法,但进展受限于缺乏通用的解释评估方法。本文探讨核心问题:什么是好解释?提出一种多元化的解释美德框架,融合科学哲学中的贝叶斯、库恩、德沃夏克和诺摩逻辑四种视角,系统评估并改进MI中的解释。研究发现,紧凑证明法综合了多种解释美德,是前景广阔的路径。框架暗示的关键研究方向包括:(1)明确定义解释简洁性,(2)聚焦统一性解释,(3)推导适用于神经网络的普适原理。更优的MI方法将提升我们对AI系统的监测、预测与引导能力。
原文摘要 · Abstract (English)
Mechanistic Interpretability (MI) aims to understand neural networks through causal explanations. Though MI has many explanation-generating methods, progress has been limited by the lack of a universal approach to evaluating explanations. Here we analyse the fundamental question "What makes a good explanation?" We introduce a pluralist Explanatory Virtues Framework drawing on four perspectives from the Philosophy of Science - the Bayesian, Kuhnian, Deutschian, and Nomological - to systematically evaluate and improve explanations in MI. We find that Compact Proofs consider many explanatory virtues and are hence a promising approach. Fruitful research directions implied by our framework include (1) clearly defining explanatory simplicity, (2) focusing on unifying explanations and (3) deriving universal principles for neural networks. Improved MI methods enhance our ability to monitor, predict, and steer AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。