用可解释概念分析神经网络表征,找关键神经元。
ConceptTracer: Interactive Analysis of Concept Saliency and Selectivity in Neural Representations
- 通过信息论度量概念显著性与选择性,定位响应特定概念的神经元。
- 在TabPFN模型上验证,能有效发现可解释的神经元激活模式。
- 适合关注模型可解释性的研究者与工程师使用。
神经网络在各类任务中表现优异,但其决策过程往往不透明。尽管对机制可解释性兴趣日益增长,针对通用神经网络表示,尤其是表格基础模型的系统性分析工具仍有限。本文提出ConceptTracer,一个交互式应用,通过人类可理解的概念视角分析神经表示。该工具整合了两种信息论度量,分别量化概念显著性与选择性,使研究人员能识别对特定概念反应强烈的神经元。我们在TabPFN学习的表示上展示了ConceptTracer的效用,结果表明该方法有助于发现可解释的神经元。这些能力共同构成一个实用框架,用于探究如TabPFN等模型如何编码概念级信息。ConceptTracer开源可用:https://github.com/ml-lab-htw/concept-tracer。
原文摘要 · Abstract (English)
Neural networks deliver impressive predictive performance across a variety of tasks, but they are often opaque in their decision-making processes. Despite a growing interest in mechanistic interpretability, tools for systematically exploring the representations learned by neural networks in general, and tabular foundation models in particular, remain limited. In this work, we introduce ConceptTracer, an interactive application for analyzing neural representations through the lens of human-interpretable concepts. ConceptTracer integrates two information-theoretic measures that quantify concept saliency and selectivity, enabling researchers and practitioners to identify neurons that respond strongly to individual concepts. We demonstrate the utility of ConceptTracer on representations learned by TabPFN and show that our approach facilitates the discovery of interpretable neurons. Together, these capabilities provide a practical framework for investigating how neural networks like TabPFN encode concept-level information. ConceptTracer is available at https://github.com/ml-lab-htw/concept-tracer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。