arXiv:2604.18712cs.CL2026-04ACL

用眼动数据验证大模型是否模拟人类阅读时长,发现早期层表现更优。

Probing for Reading Times

  • 在五种语言的语料上用线性回归分析模型各层表示与阅读时间的关系。
  • 早期层表示比困惑度更能预测首次注视和注视持续时间等早期阅读指标。
  • 不同语言和阅读指标下最佳预测器不同,混合使用效果更佳。

探针研究显示语言模型表征包含丰富的语言信息,但其是否也捕捉到人类认知处理信号尚不明确。本文通过在涵盖五种语言(英语、希腊语、希伯来语、俄语、土耳其语)的眼动数据集上使用正则化线性回归,将模型各层表征与困惑度、信息量、logit-lens困惑度等标量预测因子进行比较。结果表明,早期层表征在预测首次注视、注视持续时间等早期阅读指标时优于困惑度;而对总阅读时间这类晚期指标,尽管表征更压缩,困惑度仍表现更优。此外,在某些任务中结合困惑度与早期层表征可进一步提升性能。整体而言,最优预测器显著依赖于语言类型与具体眼动测量指标。

原文摘要 · Abstract (English)

Probing has shown that language model representations encode rich linguistic information, but it remains unclear whether they also capture cognitive signals about human processing. In this work, we probe language model representations for human reading times. Using regularized linear regression on two eye-tracking corpora spanning five languages (English, Greek, Hebrew, Russian, and Turkish), we compare the representations from every model layer against scalar predictors -- surprisal, information value, and logit-lens surprisal. We find that the representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration. The concentration of predictive power in the early layers suggests that human-like processing signatures are captured by low-level structural or lexical representations, pointing to a functional alignment between model depth and the temporal stages of human reading. In contrast, for late-pass measures such as total reading time, scalar surprisal remains superior, despite its being a much more compressed representation. We also observe performance gains when using both surprisal and early-layer representations. Overall, we find that the best-performing predictor varies strongly depending on the language and eye-tracking measure.

语言模型眼动追踪认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。