arXiv:2411.02280cs.CLcs.LG2024-11NAACL被引 42

发现大模型中类似人类语言脑区的特化单元

The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units

  • 用神经科学方法定位大模型中的语言特化单元
  • 剔除这些单元会严重损害语言任务表现
  • 发现部分模型在推理与社交能力上也有特化网络

大型语言模型(LLMs)不仅在语言任务上表现出色,还能处理逻辑推理和社会推断等非语言任务。人类大脑中已识别出一个选择性且因果性支持语言处理的核心语言系统。本文探讨大模型中是否存在类似的语言特化现象。通过采用与神经科学相同的定位方法,在18个主流大模型中识别出语言选择性单元,并通过消融实验验证其因果作用:仅剔除语言选择性单元会导致语言任务性能显著下降,而随机单元则无此效应。此外,语言选择性单元与人类语言脑区记录的对齐程度高于随机单元。进一步研究发现,部分模型在推理和社交能力方面也存在专门化网络,但不同模型间差异明显。该研究为大模型的功能特化提供了功能性和因果性证据,并揭示了其与人脑功能组织的相似性。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit remarkable capabilities on not just language tasks, but also various tasks that are not linguistic in nature, such as logical reasoning and social inference. In the human brain, neuroscience has identified a core language system that selectively and causally supports language processing. We here ask whether similar specialization for language emerges in LLMs. We identify language-selective units within 18 popular LLMs, using the same localization approach that is used in neuroscience. We then establish the causal role of these units by demonstrating that ablating LLM language-selective units -- but not random units -- leads to drastic deficits in language tasks. Correspondingly, language-selective LLM units are more aligned to brain recordings from the human language system than random units. Finally, we investigate whether our localization method extends to other cognitive domains: while we find specialized networks in some LLMs for reasoning and social capabilities, there are substantial differences among models. These findings provide functional and causal evidence for specialization in large language models, and highlight parallels with the functional organization in the brain.

大模型机制语言特化因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。