arXiv:2505.12196cs.CL2025-05被引 4

更大语言模型向量预测人类阅读时间能力反而下降,说明模型复杂度并非越高越好。

Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled

  • 用未训练的大模型控制维度膨胀,检验不同规模语言模型向量预测效果。
  • 当参数量超几十亿后,模型预测人类阅读时间能力反降,且训练优势消失。
  • 结果质疑大模型作为人脑语言处理模型的普适性,适合认知科学与模型可解释性研究者。

大型语言模型(LLM)在语言理解上的表现使其被视为人类句子处理的潜在模型,有人认为模型性能与预测能力呈正相关,即模型越强,对人类阅读时间等心理数据的拟合越好。然而近期研究表明,当使用信息论意义上的意外度(surprisal)作为预测变量时,这种增长趋势会在模型过大时逆转。另有研究使用完整向量进行预测,仍发现正向增长,但未控制维度随模型规模膨胀的问题,尤其缺乏对超过16亿参数的未训练模型的对比。本研究通过使用最多达660亿参数的未训练模型进行维度控制,发现逆向缩放现象依然存在;此外,在多数数据集上,训练过的模型相比对应未训练模型的提升在数亿参数量级后趋于零。

原文摘要 · Abstract (English)

The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive 'quality-power' relationship, in which language models' (LMs') fit to psychometric data continues to improve as their ability to predict words in context increases. This is important because it might suggest that elements of LLM architecture reflect the architecture of the human sentence processing faculty, and that any inadequacies in predicting human reading time and brain imaging data may be attributed to insufficient model complexity, which recedes as larger models become available. But recent studies have shown this scaling inverts after a point, as LMs become excessively large and accurate, when information-theoretic surprisal is used as a predictor. Other studies propose the use of entire vectors from differently sized LLMs, still showing positive scaling, casting doubt on the value of surprisal as a predictor, but do not control for dimensionality expansion using untrained LLMs with more than 1.6B parameters. This study evaluates scaling of LLM vector predictors controlled using untrained LLMs with up to 66B parameters. Results show that inverse scaling obtains, and moreover the contribution of trained LMs over corresponding untrained LMs drops to zero at around a few billion parameters on most datasets.

语言模型认知科学逆向缩放向量预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。