arXiv:2505.11485cs.CLcs.AI2025-05NeurIPS

用Transformer模型模拟阅读时的眼动行为,发现仍无法完全解释人类预测机制。

Modeling cognitive processes of natural reading with transformer-based Language Models

  • 用GPT2、LLaMA等大模型分析西班牙语阅读眼动数据
  • 模型能解释部分眼动停留时间,但仍有显著差距
  • 适合研究语言认知与AI理解差异的学者参考

自然语言处理的进展催生了强大的文本生成模型。与此同时,神经科学开始利用这些模型探索语言理解中的认知过程。先前研究显示,N-gram和LSTM网络可部分解释阅读中眼动行为(特别是注视持续时间)的可预测性效应。本研究进一步评估了基于Transformer的模型(GPT2、LLaMA-7B、LLaMA2-7B),以探究其与眼动数据的关系。结果表明,这些架构在解释来自西班牙语读者(Rioplantense)记录的注视持续时间方差方面,优于早期模型。然而,与人类可预测性所捕捉的总方差相比,模型仍未能解释全部差异。这表明,尽管技术进步显著,当前最先进的语言模型在预测语言的方式上,依然与人类读者存在本质差异。

原文摘要 · Abstract (English)

Recent advances in Natural Language Processing (NLP) have led to the development of highly sophisticated language models for text generation. In parallel, neuroscience has increasingly employed these models to explore cognitive processes involved in language comprehension. Previous research has shown that models such as N-grams and LSTM networks can partially account for predictability effects in explaining eye movement behaviors, specifically Gaze Duration, during reading. In this study, we extend these findings by evaluating transformer-based models (GPT2, LLaMA-7B, and LLaMA2-7B) to further investigate this relationship. Our results indicate that these architectures outperform earlier models in explaining the variance in Gaze Durations recorded from Rioplantense Spanish readers. However, similar to previous studies, these models still fail to account for the entirety of the variance captured by human predictability. These findings suggest that, despite their advancements, state-of-the-art language models continue to predict language in ways that differ from human readers.

语言模型认知科学眼动研究Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。