arXiv:2604.15503cs.CL2026-04

不同语言训练的模型在脑科学评分上表现相似,揭示了通用结构学习能力。

Brain Score Tracks Shared Properties of Languages: Evidence from Many Natural Languages and Structured Sequences

论文配图:Brain Score Tracks Shared Properties of Languages: Evidence from Many Natural Languages and Structured Sequences
图 1 · 摘自论文原文
  • 用脑科学评分比较多种语言和结构化数据训练的模型表现
  • 跨语言模型的评分相近,结构化数据模型也表现不俗
  • 提示该评分可能反映通用结构识别而非人类语言处理机制

近期基于神经网络的语言模型突破引发了关于其处理方式是否接近人类语言处理的疑问。采用脑科学评分(Brain Score, BS)框架——从语言模型激活预测阅读时的fMRI信号——已有研究支持两者存在高度相似性。为理解这一相似性,我们训练了多种输入数据上的语言模型并评估其在BS上的表现。结果发现,使用多种语系自然语言训练的模型在BS上表现极为接近;而基于人类基因组、Python代码及纯层次结构(嵌套括号)训练的模型也表现出合理水平,部分情况下与自然语言模型相当。这些发现表明,脑科学评分能揭示语言模型提取跨自然语言共通结构的能力,但该指标可能不足以仅凭高分就推断出模型具备类人语言处理特性。

原文摘要 · Abstract (English)

Recent breakthroughs in language models (LMs) using neural networks have raised the question: how similar are these models' processing to human language processing? Results using a framework called Brain Score (BS) -- predicting fMRI activations during reading from LM activations -- have been used to argue for a high degree of similarity. To understand this similarity, we conduct experiments by training LMs on various types of input data and evaluate them on BS. We find that models trained on various natural languages from many different language families have very similar BS performance. LMs trained on other structured data -- the human genome, Python, and pure hierarchical structure (nested parentheses) -- also perform reasonably well and close to natural languages in some cases. These findings suggest that BS can highlight language models' ability to extract common structure across natural languages, but that the metric may not be sensitive enough to allow us to infer human-like processing from a high BS score alone.

语言模型脑科学评分结构学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。