首次为神经元解释提供理论保障,确保结果既准确又稳定。
Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
- 将神经元识别视为机器学习的逆过程,推导出可信赖解释的理论基础。
- 提出置信区间方法,使解释结果在不同数据集上保持一致,覆盖概率可保证。
- 适用于需要可信机制解释的AI安全、医疗等高风险场景。
神经元识别是机制可解释性中的常用工具,旨在揭示深度网络中单个神经元所表征的人类可理解概念。尽管如Network Dissection和CLIP-Dissect等算法取得了显著的实证成功,但缺乏严格的理论基础,这制约了其可信度与可靠性。本文观察到神经元识别可被视为机器学习的逆过程,由此推导出对解释结果的理论保证。基于此,我们首次对两个核心挑战进行理论分析:(1) 忠实性——识别出的概念是否真实反映神经元的底层功能;(2) 稳定性——识别结果在不同探测数据集上是否一致。我们推导了广泛使用的相似性度量(如准确率、AUROC、IoU)的泛化界,以保证忠实性;并提出一种自助集成方法(bootstrap ensemble),结合BE(Bootstrap Explanation)方法,生成具有保证覆盖率的概率概念预测集。在合成数据和真实数据上的实验验证了理论结果,并展示了方法的实用性,为可信神经元识别迈出了关键一步。
原文摘要 · Abstract (English)
Neuron identification is a popular tool in mechanistic interpretability, aiming to uncover the human-interpretable concepts represented by individual neurons in deep networks. While algorithms such as Network Dissection and CLIP-Dissect achieve great empirical success, a rigorous theoretical foundation remains absent, which is crucial to enable trustworthy and reliable explanations. In this work, we observe that neuron identification can be viewed as the inverse process of machine learning, which allows us to derive guarantees for neuron explanations. Based on this insight, we present the first theoretical analysis of two fundamental challenges: (1) Faithfulness: whether the identified concept faithfully represents the neuron's underlying function and (2) Stability: whether the identification results are consistent across probing datasets. We derive generalization bounds for widely used similarity metrics (e.g. accuracy, AUROC, IoU) to guarantee faithfulness, and propose a bootstrap ensemble procedure that quantifies stability along with BE (Bootstrap Explanation) method to generate concept prediction sets with guaranteed coverage probability. Experiments on both synthetic and real data validate our theoretical results and demonstrate the practicality of our method, providing an important step toward trustworthy neuron identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。