不同语法一致现象共用大模型中的相同神经单元。
Different types of syntactic agreement recruit the same units within large language models
- 通过类认知神经科学方法定位模型中响应语法现象的神经单元。
- 不同语法一致类型共享大量神经单元,跨语言分析也支持此规律。
- 适合对语言模型内部表征机制感兴趣的学者研究。
大型语言模型(LLMs)能可靠区分语法正确与错误的句子,但其内部如何表征语法知识仍不明确。我们采用受认知神经科学启发的功能定位方法,识别了七种开源模型中对67种英语语法现象最敏感的神经单元。这些单元在包含相应现象的句子中被持续激活,并且因果上支持模型的语法表现能力。关键发现是:不同类型的语法一致(如主谓、代词、限定词-名词)招募重叠的神经单元集合,表明一致性在模型中构成一个有意义的功能类别。该模式在英语、俄语和汉语中均成立;进一步跨语言分析显示,在57种不同语言中,结构更相似的语言在主谓一致上共享更多神经单元。综合来看,这些结果揭示语法一致——这一语法依赖的关键标志——构成了大语言模型表征空间中的有意义类别。
原文摘要 · Abstract (English)
Large language models (LLMs) can reliably distinguish grammatical from ungrammatical sentences, but how grammatical knowledge is represented within the models remains an open question. We investigate whether different syntactic phenomena recruit shared or distinct components in LLMs. Using a functional localization approach inspired by cognitive neuroscience, we identify the LLM units most responsive to 67 English syntactic phenomena in seven open-weight models. These units are consistently recruited across sentences containing the phenomena and causally support the models' syntactic performance. Critically, different types of syntactic agreement (e.g., subject-verb, anaphor, determiner-noun) recruit overlapping sets of units, suggesting that agreement constitutes a meaningful functional category for LLMs. This pattern holds in English, Russian, and Chinese; and further, in a cross-lingual analysis of 57 diverse languages, structurally more similar languages share more units for subject-verb agreement. Taken together, these findings reveal that syntactic agreement-a critical marker of syntactic dependencies-constitutes a meaningful category within LLMs' representational spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。