arXiv:2604.01404cs.CLcs.AI2026-04被引 1

发现语言模型中能精准定位实体的稀疏神经元,可直接操控事实记忆。

Friends and Grandmothers in Silico: Localizing Entity Cells in Language Models

  • 通过激活一致性筛选出对特定实体敏感的MLP神经元(称作实体细胞)
  • 抑制单个实体细胞会精准删除对应实体知识,激活它即可恢复知识
  • 该机制跨拼写、多语言和指令微调保持稳定,适合可解释性研究

语言模型如何从参数中检索特定实体的事实?我们通过搜索稀疏且实体特异的MLP神经元——类比神经科学中的'祖母细胞'假说——来探究这一问题,并测试它们在事实回忆中的因果作用。通过在七种模型上对PopQA的精选子集进行分析,依据不同提示下同一实体的激活一致性排名MLP神经元,定位候选实体细胞。所有模型中,这些神经元主要集中在早期层,这一现象并非架构强制所致。以Qwen2.5-7B base为实验对象,发现最明确的因果证据:抑制一个定位到的细胞会精确消除其对应实体的回忆,而其他实体不受影响;仅激活一个细胞即可恢复大多数实体的知识,即使实体未出现在上下文中。相同细胞在别名、缩写、拼写错误及多语言形式下仍被识别,且在指令微调后保持稳定,表明其编码的是实体的规范身份而非表面标记模式。因果信号在不同模型家族间存在差异,暗示了实体知识组织方式的架构差异。这些发现为理解、控制和纠正语言模型的事实知识提供了具体可解释的切入点,并与神经科学中概念稀疏编码的长期问题形成意外的实证呼应。

原文摘要 · Abstract (English)

How do language models retrieve entity-specific facts from their parameters? We investigate this question by searching for sparse, entity-selective MLP neurons - which we call entity cells, by analogy to the "grandmother cell" hypothesis in neuroscience - and testing whether they play a causal role in factual recall. We localize candidate entity cells by ranking MLP neurons for activation consistency across varied prompts about the same entity, applying this procedure across seven models on a curated subset of PopQA. In all models, localized neurons cluster predominantly in early layers, an empirical pattern not imposed by the architecture. Using Qwen2.5-7B base as a model organism, we find the clearest causal evidence: suppressing a localized cell selectively erases recall for its matched entity while leaving others intact, and activating a single cell is sufficient to recover correct knowledge for most entities - even when the entity is absent from the context. The same cells are recovered under aliases, acronyms, misspellings, and multilingual surface forms, and remain stable through instruction tuning, suggesting they encode canonical entity identity rather than surface token patterns. Causal signals vary across model families, pointing to architectural differences in how entity knowledge is organized. These findings offer concrete, interpretable access points for understanding, controlling, and correcting factual knowledge in language models, and draw a surprising empirical parallel to longstanding questions in neuroscience about sparse coding of concepts.

可解释性实体记忆神经元分析因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。