arXiv:2607.20995cs.CL2026-07

发现大模型处理生物性概念的内部神经机制,揭示其分布且依赖上下文的特性。

Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept

论文配图:Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept
图 1 · 摘自论文原文
  • 通过构造最小对比数据集,定位大模型中负责生物性判断的神经通路。
  • 找到一个因果相关但分布式的生物性神经电路,跨模型泛化能力有限。
  • 该研究为理解模型如何捕捉复杂语义提供新方法,适合关注模型可解释性的研究者。

在书面语言中区分有生命与无生命的概念,不仅需要浅层文本处理,还需识别复杂的选言约束和上下文线索,如动词-论元互动。然而,当前大型语言模型(LLMs)似乎具备这种能力。我们探究这种对生物性敏感的行为是否可追溯到一组局部化的因果相关组件与连接。为此,我们构建了一个受控的最小对比数据集,并对四款开源权重模型进行电路发现。通过深入实验与消融分析,我们证实存在一个负责处理生物性的因果机制,从而发现了生物性电路。同时,该电路的局部性较弱,且在不同模型及生物性任务间仅部分泛化,验证了生物性概念的分布式、上下文依赖性和一定程度的渐进性特征。

原文摘要 · Abstract (English)

Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.

模型可解释性语义理解神经电路大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。