arXiv:2506.05774cs.LG2025-06ICML被引 11

统一神经元解释评估框架,发现多数指标不可靠。

Evaluating Neuron Explanations: A Unified Framework with Sanity Checks

  • 构建统一数学框架,整合多种解释评估方法。
  • 提出两项合理性检验,发现多数指标对标签改动无反应。
  • 给出可靠评估指标清单与未来评估指南。

理解神经网络中单个单元的功能是机制可解释性的关键。通常通过生成简单的文本说明来解释神经元行为。这些解释是否可信,取决于其评估方法的可靠性。本文将多种现有解释评估方法统一到一个数学框架下,使评估流程更清晰,并可应用统计方法。此外,提出两项简单的合理性检验,发现许多常用指标在概念标签被大幅更改后仍保持原分值,表明其不可靠。基于实验与理论分析,本文提出未来评估应遵循的准则,并识别出一组可靠的评估指标。

原文摘要 · Abstract (English)

Understanding the function of individual units in a neural network is an important building block for mechanistic interpretability. This is often done by generating a simple text explanation of the behavior of individual neurons or units. For these explanations to be useful, we must understand how reliable and truthful they are. In this work we unify many existing explanation evaluation methods under one mathematical framework. This allows us to compare existing evaluation metrics, understand the evaluation pipeline with increased clarity and apply existing statistical methods on the evaluation. In addition, we propose two simple sanity checks on the evaluation metrics and show that many commonly used metrics fail these tests and do not change their score after massive changes to the concept labels. Based on our experimental and theoretical results, we propose guidelines that future evaluations should follow and identify a set of reliable evaluation metrics.

神经元解释评估框架可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。