简单n-gram模型比复杂Transformer更准预测阅读时长。
N-gram-like Language Models Predict Naturalistic Reading Time Best
- 用n-gram统计解释阅读时间,而非复杂模型的深层特征
- n-gram概率与眼动追踪阅读时长相关性最高
- 适合关注阅读行为建模与认知计算的研究者
近期研究发现,当代如Transformer等语言模型在下一个词预测上表现优异,但其生成的概率反而降低了对自然文本阅读时长的预测能力。本文提出,阅读时长主要受简单n-gram统计规律支配,而非先进Transformer模型所学习的复杂统计特征。我们证明,那些预测结果与n-gram概率最相关的神经语言模型,其概率计算也与自然文本上的眼动追踪阅读时长指标相关性最强。
原文摘要 · Abstract (English)
Recent work has found that contemporary language models such as transformers can become so good at next-word prediction that the probabilities they calculate become worse for predicting naturalistic reading time. In this paper, we propose that this can be explained by reading time being shaped by simple n-gram statistics rather than the more complex statistics learned by state-of-the-art transformer language models. We demonstrate that the neural language models whose predictions are most correlated with n-gram probability are also those that calculate probabilities that are the most correlated with eye-tracking-based metrics of reading time on naturalistic text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。