arXiv:2603.24953cs.CV2026-03中稿 · CVPR

提出验证神经元概念的新框架,提升解释准确性

Select, Hypothesize and Verify: Towards Verified Neuron Concept Interpretation

  • 通过激活分布分析筛选典型样本,再生成并验证概念
  • 新方法使概念激活神经元的概率高出现有方法约1.5倍
  • 适合需要可信神经网络解释的研究者使用

理解神经网络决策的关键在于解释神经元的功能(即概念)。现有方法通过生成自然语言描述来解释神经元概念,推动了对模型决策机制的理解。然而,这些方法假设每个神经元都有明确功能且提供判别性特征,实际上部分神经元可能冗余或产生误导性概念,导致对决策因素的误读。为此,本文提出神经元功能验证机制,检验生成的概念是否显著激活对应神经元。我们构建了「选择-假设-验证」框架:首先通过激活分布分析筛选能最好体现神经元功能行为的激活样本;其次对选定神经元形成概念假设;最后验证生成概念是否真实反映神经元功能。大量实验表明,该方法生成的概念更准确,其激活对应神经元的概率约为当前最优方法的1.5倍。

原文摘要 · Abstract (English)

It is essential for understanding neural network decisions to interpret the functionality (also known as concepts) of neurons. Existing approaches describe neuron concepts by generating natural language descriptions, thereby advancing the understanding of the neural network's decision-making mechanism. However, these approaches assume that each neuron has well-defined functions and provides discriminative features for neural network decision-making. In fact, some neurons may be redundant or may offer misleading concepts. Thus, the descriptions for such neurons may cause misinterpretations of the factors driving the neural network's decisions. To address the issue, we introduce a verification of neuron functions, which checks whether the generated concept highly activates the corresponding neuron. Furthermore, we propose a Select-Hypothesize-Verify framework for interpreting neuron functionality. This framework consists of: 1) selecting activation samples that best capture a neuron's well-defined functional behavior through activation-distribution analysis; 2) forming hypotheses about concepts for the selected neurons; and 3) verifying whether the generated concepts accurately reflect the functionality of the neuron. Extensive experiments show that our method produces more accurate neuron concepts. Our generated concepts activate the corresponding neurons with a probability approximately 1.5 times that of the current state-of-the-art method.

神经元解释概念验证可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。