研究发现语言模型的意外度预测在阅读中不总有效,尤其在重复阅读时失效。
The Effect of Surprisal on Reading Times in Information Seeking and Repeated Reading
- 用眼动数据测试信息搜索和重复阅读中的语言意外度影响。
- 标准意外度估计仍能预测阅读时间,但任务特定估计效果差。
- 提示当前语言模型与人类认知存在偏差,不适合直接用于心理预测。
surprisal 对处理难度的影响是心理语言学的核心议题。本文利用眼动追踪数据,考察日常生活中常见的三种语言处理情境:信息搜索、重复处理及二者结合。使用通用的意外度估算方法,发现预期的线性关系依然成立。然而,当采用与人类任务和情境匹配的特定场景意外度估计时,信息搜索中其预测能力并未优于标准估算;而在重复阅读中,特定情境下的意外度接近零,对阅读时间无预测力。这些结果揭示了人类任务与记忆表征同当前语言模型之间的错配,质疑了此类模型在估算认知相关量时的可靠性。文中进一步讨论了由此带来的理论挑战。
原文摘要 · Abstract (English)
The effect of surprisal on processing difficulty has been a central topic of investigation in psycholinguistics. Here, we use eyetracking data to examine three language processing regimes that are common in daily life but have not been addressed with respect to this question: information seeking, repeated processing, and the combination of the two. Using standard regime-agnostic surprisal estimates we find that the prediction of surprisal theory regarding the presence of a linear effect of surprisal on processing times, extends to these regimes. However, when using surprisal estimates from regime-specific contexts that match the contexts and tasks given to humans, we find that in information seeking, such estimates do not improve the predictive power of processing times compared to standard surprisals. Further, regime-specific contexts yield near zero surprisal estimates with no predictive power for processing times in repeated reading. These findings point to misalignments of task and memory representations between humans and current language models, and question the extent to which such models can be used for estimating cognitively relevant quantities. We further discuss theoretical challenges posed by these results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。