arXiv:2603.23624cs.CL2026-03

人类阅读中不存在实时挖坑效应,反而支持神经语言模型的预测。

Revisiting Real-Time Digging-In Effects: No Evidence from NP/Z Garden-Paths

  • 通过自定速阅读和Maze实验,对比人类与大模型对复杂句式的反应。
  • 非句末歧义项显示反向挖坑趋势,与神经模型预测一致。
  • 句末歧义项的正向趋势可能源于结尾处理干扰,非真实认知过程。

挖坑效应指句中歧义区域越长,解析难度越高,被视作自组织句法处理的证据。但意外理论预测若长度未改变统计预期,则不应出现该效应,且神经语言模型表现出相反模式。当前尚不清楚该效应是否为人类实时句法处理的可靠现象,抑或源于结尾处理或方法混杂。本研究通过两个实验,在英语名词短语/零代词花园路径句上使用Maze和自定速阅读,比较人类行为与多个大型语言模型的预测。结果未发现实时挖坑效应。关键发现:句末与非句末歧义项呈现不同模式——仅在句末出现正向挖坑趋势,可能受收尾过程干扰;而非句末项(更纯净的实时处理测试)则显示反向趋势,与神经模型预测一致。

原文摘要 · Abstract (English)

Digging-in effects, where disambiguation difficulty increases with longer ambiguous regions, have been cited as evidence for self-organized sentence processing, in which structural commitments strengthen over time. In contrast, surprisal theory predicts no such effect unless lengthening genuinely shifts statistical expectations, and neural language models appear to show the opposite pattern. Whether digging-in is a robust real-time phenomenon in human sentence processing -- or an artifact of wrap-up processes or methodological confounds -- remains unclear. We report two experiments on English NP/Z garden-path sentences using Maze and self-paced reading, comparing human behavior with predictions from an ensemble of large language models. We find no evidence for real-time digging-in effects. Critically, items with sentence-final versus nonfinal disambiguation show qualitatively different patterns: positive digging-in trends appear only sentence-finally, where wrap-up effects confound interpretation. Nonfinal items -- the cleaner test of real-time processing -- show reverse trends consistent with neural model predictions.

句法处理认知语言学神经语言模型心理实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。