arXiv:2506.01172cs.CL2025-06ACL被引 5

大模型预测阅读时间变差,不是因为训练数据泄露。

The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage

  • 检查五大数据集的文本重叠度,发现泄露极轻微。
  • 即使在无泄露数据上训练,大模型仍表现更差。
  • 适合语言认知与模型可解释性研究者阅读。

在心理语言学建模中,更大的预训练语言模型产生的突兀度(surprisal)反而更难预测自然语境下的人类阅读时长。有观点认为这可能源于模型在训练时接触过测试文本导致的数据泄露。本文通过两项大规模研究检验这一假设:第一项分析显示,五个自然阅读时长语料库与两个预训练数据集在词元n-gram长度和频率上的重叠极少;第二项研究使用几乎无重叠的‘无泄露’数据训练模型,依然复现了模型规模越大、突兀度与阅读时长匹配度越低的结果。这表明先前结论并非由数据泄露引起。

原文摘要 · Abstract (English)

In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been speculated that this may be due to data leakage that caused language models to see the text stimuli during training. This paper presents two studies to address this concern at scale. The first study reveals relatively little leakage of five naturalistic reading time corpora in two pre-training datasets in terms of length and frequency of token $n$-gram overlap. The second study replicates the negative relationship between language model size and the fit of surprisal to reading times using models trained on 'leakage-free' data that overlaps only minimally with the reading time corpora. Taken together, this suggests that previous results using language models trained on these corpora are not driven by the effects of data leakage.

语言模型认知科学数据泄露突兀度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。