arXiv:2505.14424cs.LG2025-05

用逻辑与贝叶斯结合的方法解释神经网络,让每个神经元说出它支持的判断理由。

Explaining Neural Networks with Reasons

  • 为每个神经元生成'理由向量',量化其对特定命题的支持度。
  • 方法在图像分类、文本情感等任务中验证有效,能准确反映模型决策依据。
  • 适合需要可解释性保障的场景,如医疗、金融等高风险领域。

我们提出一种新的神经网络可解释性方法,基于新颖的数学哲学理论中的'理由'概念。该方法为每个神经元计算一个称为'理由向量'的向量,进而评估该向量对各类命题(如输入图像表示数字2或输入提示具有负面情感)的支持强度。这实现了对神经元及其组合的解释,融合了逻辑与贝叶斯视角,并能处理多语义性问题(即单个神经元可参与多个概念)。我们从理论和实证上证明该方法:(1) 建立在哲学公认的解释概念之上;(2) 具有统一性,适用于主流神经网络架构与模态;(3) 可扩展,仅需前向传播即可计算理由向量;(4) 忠实性高,基于理由向量干预神经元会引发预期输出变化;(5) 正确性好,模型理由结构与数据源一致;(6) 可训练,可通过训练优化理由强度;(7) 实用性强,有助于提升模型鲁棒性与公平性。

原文摘要 · Abstract (English)

We propose a new interpretability method for neural networks, which is based on a novel mathematico-philosophical theory of reasons. Our method computes a vector for each neuron, called its reasons vector. We then can compute how strongly this reasons vector speaks for various propositions, e.g., the proposition that the input image depicts digit 2 or that the input prompt has a negative sentiment. This yields an interpretation of neurons, and groups thereof, that combines a logical and a Bayesian perspective, and accounts for polysemanticity (i.e., that a single neuron can figure in multiple concepts). We show, both theoretically and empirically, that this method is: (1) grounded in a philosophically established notion of explanation, (2) uniform, i.e., applies to the common neural network architectures and modalities, (3) scalable, since computing reason vectors only involves forward-passes in the neural network, (4) faithful, i.e., intervening on a neuron based on its reason vector leads to expected changes in model output, (5) correct in that the model's reasons structure matches that of the data source, (6) trainable, i.e., neural networks can be trained to improve their reason strengths, (7) useful, i.e., it delivers on the needs for interpretability by increasing, e.g., robustness and fairness.

神经网络解释可解释性理由向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。