大模型语言能力发展与人脑语言区的匹配度,主要取决于语法规则而非常识推理。
From Language to Cognition: How LLMs Outgrow the Human Language Network
- 通过34个训练节点分析,发现脑对齐更贴近语法知识发展
- 模型超越人类语言水平后,脑对齐与预测能力相关性下降
- 适合研究语言神经机制或模型认知本质的研究者
大型语言模型(LLMs)表现出与人类语言网络神经活动的高度相似性。然而,语言如何塑造类脑表征,以及这种表征在不同任务下随训练演化的规律仍不清晰。本文针对8种不同规模、共300B token训练过程中的34个检查点进行基准测试,分析脑对齐与语言能力的关系。结果表明,脑对齐更紧密地追踪形式语言能力(即语法规则知识)的发展,而非功能性语言能力(包含世界知识和推理)。尽管功能性能力持续提升,但其与脑对齐的关联较弱,说明人类语言网络主要编码形式语言结构而非广泛认知功能。进一步发现,在控制特征大小后,模型规模并非脑对齐的可靠预测因子;当模型超过人类语言水平后,下一个词预测、行为对齐与脑对齐的相关性减弱。基于迄今最全面的语言神经基准集,我们发现语言脑对齐指标尚未饱和,提示未来模型仍有优化空间。综合来看,人类语言网络更适合用语言的形式层面而非功能层面来建模。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language shaping brain-like representations, and their evolution during training as a function of different tasks remain unclear. We here benchmark 34 training checkpoints spanning 300B tokens across 8 different model sizes to analyze how brain alignment relates to linguistic competence. Specifically, we find that brain alignment tracks the development of formal linguistic competence -- i.e., knowledge of linguistic rules -- more closely than functional linguistic competence. While functional competence, which involves world knowledge and reasoning, continues to develop throughout training, its relationship with brain alignment is weaker, suggesting that the human language network primarily encodes formal linguistic structure rather than broader cognitive functions. We further show that model size is not a reliable predictor of brain alignment when controlling for feature size and find that the correlation between next-word prediction, behavioral alignment and brain alignment fades once models surpass human language proficiency. Finally, using the largest set of rigorous neural language benchmarks to date, we show that language brain alignment benchmarks remain unsaturated, highlighting opportunities for improving future models. Taken together, our findings suggest that the human language network is best modeled by formal, rather than functional, aspects of language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。