arXiv:2508.05803cs.CL2025-08Transactions of th…被引 4

给Transformer模型加短暂记忆,能提升语言学习但降低阅读时间预测能力

Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models

  • 在训练中引入人类短暂记忆机制,让模型遗忘旧词形
  • 语言建模性能和句法理解能力均显著提升
  • 适合研究记忆机制对学习的影响,不适用于行为预测

人类记忆具有短暂性,输入句子中的具体词形会迅速消失。认知科学认为这种记忆限制反而有助于语言学习,经典连接主义模型也支持此观点。尽管变压器(Transformers)模型在无记忆限制的情况下仍能有效学习语言,我们通过严格控制实验,在发展性真实训练集上对比有无短暂记忆的变压器模型。结果发现,引入短暂记忆可稳定提升语言建模性能及特定句法评估表现;但出人意料的是,其基于意外度(surprisal)的人类阅读时间预测能力反而下降。进一步分析表明,这一矛盾无法用已有解释(如优秀模型更难拟合人类阅读时间)说明。结果支持记忆限制对神经网络语言学习有益,但对行为预测无益。

原文摘要 · Abstract (English)

Human memory is fleeting. As words are processed, the exact wordforms that make up incoming sentences are rapidly lost. Cognitive scientists have long believed that this limitation of memory may, paradoxically, help in learning language - an idea supported by classic connectionist modelling work. The rise of Transformers appears to challenge this idea, as these models can learn language effectively, despite lacking memory limitations or other architectural recency biases. Here, we investigate the hypothesized benefit of fleeting memory for language learning in tightly controlled experiments on transformer language models. Training transformers with and without fleeting memory on a developmentally realistic training set, we find that fleeting memory consistently improves language learning (as quantified by both overall language modelling performance and targeted syntactic evaluation) but, unexpectedly, impairs surprisal-based prediction of human reading times. Interestingly, follow up analyses revealed that this discrepancy - better language modeling, yet worse reading time prediction - could not be accounted for by prior explanations of why better language models sometimes fit human reading time worse. Together, these results support a benefit of memory limitations on neural network language learning - but not on predicting behavior.

语言模型记忆机制阅读时间预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。