arXiv:2605.15440cs.CL2026-05被引 1

语言模型预测人类阅读难度时偏差大,因解析能力更强。

Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis

论文配图:Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
图 1 · 摘自论文原文
  • 用神经网络语法模型控制同时解析数,模拟不同认知负荷
  • 减少并行解析数后,模型预测的歧义句困难度上升但仍不足
  • 说明语言模型与人类在句子解析能力上存在根本差异

surprisal 理论认为词语处理难度由其上下文可预测性决定,为语言模型的下一个词预测与人类句子处理之间提供了可能联系。尽管语言模型的困惑度能有效预测自然文本中的阅读时间,但在受控的句法歧义研究中(特别是花园路径句)系统性低估了实际困难程度。这可能源于人类与语言模型在计算约束上的差异。本文检验一种假设:语言模型可能比人类能同时考虑更多不同的句子解释。通过使用具有词同步束搜索的循环神经网络语法(RNNG),我们系统改变用于计算词语困惑度的并行解析数量,并用这些困惑度预测人类阅读时间。减少并行活跃解析数量确实增加了对花园路径效应的预测强度,但仍未达到人类观察到的实际效应幅度。这表明,语言模型与人类在可同时处理的解析数量上的差异无法弥合模型预测与人类句法处理之间的差距。

原文摘要 · Abstract (English)

Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and next-word predictions from language models. While language model (LM) surprisals successfully predict reading times in naturalistic text, they systematically underpredict the magnitude of difficulty observed in controlled studies of syntactic ambiguity, particularly in garden path sentences. This mismatch might arise from differences in the computational constraints between humans and LMs. Here we test one such hypothesis, specifically, that LMs may be able to simultaneously consider a greater number of distinct sentence interpretations at once, compared to humans. Using Recurrent Neural Network Grammars (RNNGs) with word-synchronous beam search, we systematically vary the number of simultaneous parses used to compute word surprisal, and then use these surprisals to predict human reading times. Reducing the number of simultaneous active parses indeed increases the magnitude of predicted garden path effects, but not nearly enough to capture the full magnitude of the effects in humans. This suggests that differences in the number of simultaneous parses available to LMs and humans cannot reconcile LM-based surprisal with human sentence processing.

语言模型句法歧义认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。