arXiv:2506.11338cs.CL2025-06Conference of the …被引 4

更大语言模型的意外度反而更难预测人脑反应

Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly

  • 用17个不同规模模型计算词级意外度
  • 大模型意外度对脑成像数据预测力更差
  • 结论不局限于阅读时长,也适用于大脑影像

近年来,基于Transformer的语言模型(LM)的意外度被广泛用于预测人类句子处理难度。已有研究发现,语言模型参数越多、训练数据越丰富,其意外度对阅读时长的预测能力反而越弱,呈现反向缩放关系。然而这些研究仅聚焦于延迟类指标。此前对脑成像数据的测试因模型数量有限,结果不明确,未能确定该现象是否仅限于行为数据。本研究通过使用来自三个不同家族的17个预训练语言模型,在两个功能磁共振成像(fMRI)数据集上进行更全面评估,结果表明:无论在哪个数据集,模型词级估计概率与其对脑活动预测效果之间的反向缩放关系依然存在。这解决了以往研究的不确定性,证明该趋势并非仅限于行为延迟指标。

原文摘要 · Abstract (English)

There has been considerable interest in using surprisal from Transformer-based language models (LMs) as predictors of human sentence processing difficulty. Recent work has observed an inverse scaling relationship between Transformers' per-word estimated probability and the predictive power of their surprisal estimates on reading times, showing that LMs with more parameters and trained on more data are less predictive of human reading times. However, these studies focused on predicting latency-based measures. Tests on brain imaging data have not shown a trend in any direction when using a relatively small set of LMs, leaving open the possibility that the inverse scaling phenomenon is constrained to latency data. This study therefore conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs across three different LM families on two functional magnetic resonance imaging (fMRI) datasets. Results show that the inverse scaling relationship between models' per-word estimated probability and model fit on both datasets still obtains, resolving the inconclusive results of previous work and indicating that this trend is not specific to latency-based measures.

语言模型脑科学意外度fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。