arXiv:2501.08618cs.CLcs.AI2025-01被引 2

大模型能自发区分语法层级与线性结构,且机制独立。

Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models

  • 用不同语言生成层级/线性语法输入,测试模型响应差异。
  • 发现处理层级语法的神经组件与处理线性语法的组件互不重叠。
  • 即使在无意义词上,层级敏感组件仍有效,说明非语义驱动。

所有自然语言均具有层次结构。在人类中,这种结构限制由神经机制编码:当两种语法使用相同词汇时,语言处理脑区仅对层次结构敏感。我们利用大规模语言模型(LLMs)探究:仅通过大规模语言分布的暴露,是否可能产生功能上不同的层次结构处理区域。我们使用英语、意大利语、日语或伪词生成输入,使底层语法符合层次或线性/位置规则。首先观察到模型在层次与线性结构输入上表现出明显不同的行为;其次发现处理层次语法的组件与处理线性语法的组件相互独立,通过消融实验验证了因果关系;最后发现,对层次结构敏感的组件在伪词语法上也活跃,表明其对层次的敏感性不依赖于语义或分布内输入。

原文摘要 · Abstract (English)

All natural languages are structured hierarchically. In humans, this structural restriction is neurologically coded: when two grammars are presented with identical vocabularies, brain areas responsible for language processing are only sensitive to hierarchical grammars. Using large language models (LLMs), we investigate whether such functionally distinct hierarchical processing regions can arise solely from exposure to large-scale language distributions. We generate inputs using English, Italian, Japanese, or nonce words, varying the underlying grammars to conform to either hierarchical or linear/positional rules. Using these grammars, we first observe that language models show distinct behaviors on hierarchical versus linearly structured inputs. Then, we find that the components responsible for processing hierarchical grammars are distinct from those that process linear grammars; we causally verify this in ablation experiments. Finally, we observe that hierarchy-selective components are also active on nonce grammars; this suggests that hierarchy sensitivity is not tied to meaning, nor in-distribution inputs.

语言模型语法解析神经机制层级结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。