arXiv:2601.03657cs.LGcs.AI2026-01被引 2

发现表格模型中存在可解释的特化神经元。

In Search of Grandmother Cells: Tracing Interpretable Neurons in Tabular Representations

  • 用信息论方法量化神经元对单一概念的敏感度和专属性。
  • 在TabPFN模型中找到对高层概念有显著响应的神经元。
  • 无需复杂技术即可识别可解释神经元,适合模型可解释性研究者。

基础模型虽强大但决策过程常不透明。神经科学与人工智能领域长期关注是否存在类似‘祖母细胞’的神经元——即仅对单一概念响应、天生可解释的神经元。本文提出两种基于信息论的度量方法,用于评估神经元对单一概念的显著性与专属性。将这些指标应用于表格基础模型TabPFN的表示层,通过简单搜索找出最显著且专一的神经元-概念配对。分析首次提供了证据:此类模型中部分神经元对高层概念表现出中等程度但统计显著的显著性与专属性。结果表明,可解释神经元可自然涌现,且在某些情况下无需依赖复杂可解释性技术即可被识别。

原文摘要 · Abstract (English)

Foundation models are powerful yet often opaque in their decision-making. A topic of continued interest in both neuroscience and artificial intelligence is whether some neurons behave like grandmother cells, i.e., neurons that are inherently interpretable because they exclusively respond to single concepts. In this work, we propose two information-theoretic measures that quantify the neuronal saliency and selectivity for single concepts. We apply these metrics to the representations of TabPFN, a tabular foundation model, and perform a simple search across neuron-concept pairs to find the most salient and selective pair. Our analysis provides the first evidence that some neurons in such models show moderate, statistically significant saliency and selectivity for high-level concepts. These findings suggest that interpretable neurons can emerge naturally and that they can, in some cases, be identified without resorting to more complex interpretability techniques.

可解释性神经元分析表格模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。