arXiv:2510.03156cs.CL2025-10

大语言模型与人脑语言响应的关联,仅在人类语言训练模型中成立。

Neural Correlates of Language Models Are Specific to Human Language

  • 通过降维和新相似性度量验证了模型与大脑的对应关系
  • 仅人类语言训练的模型与人脑活动有显著相关性
  • 位置编码是实现这种对应的关键因素

先前研究发现大型语言模型的隐藏状态与人类语言任务下的fMRI脑区响应存在相关性,被视作模型与大脑表征相似的证据。本研究检验了这些结果对多种潜在问题的稳健性:(i)降维后相关性依然存在,排除了高维诅咒的影响;(ii)采用新相似性度量后结果仍成立;(iii)相关性仅出现在人类语言训练的模型中,说明其特异性;(iv)结果依赖于模型中位置编码的存在。这些发现强化了早期研究结论,并推动关于先进语言模型生物合理性与可解释性的讨论。

原文摘要 · Abstract (English)

Previous work has shown correlations between the hidden states of large language models and fMRI brain responses, on language tasks. These correlations have been taken as evidence of the representational similarity of these models and brain states. This study tests whether these previous results are robust to several possible concerns. Specifically this study shows: (i) that the previous results are still found after dimensionality reduction, and thus are not attributable to the curse of dimensionality; (ii) that previous results are confirmed when using new measures of similarity; (iii) that correlations between brain representations and those from models are specific to models trained on human language; and (iv) that the results are dependent on the presence of positional encoding in the models. These results confirm and strengthen the results of previous research and contribute to the debate on the biological plausibility and interpretability of state-of-the-art large language models.

语言模型脑科学神经相关性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。