用混合模型解析阅读歧义,发现大模型预测力不足。
Deconstructing sentence disambiguation by joint latent modeling of reading paradigms: LLM surprisal is not enough
- 构建多范式混合潜变量模型,区分歧义概率、代价与重分析成本。
- 模型更准确还原重读行为与理解判断,优于基于GPT-2的突变值模型。
- 适合研究人类语言处理机制或模型可解释性的学者参考。
以临时歧义的花园路径句(如“当球队训练时,前锋感到疑惑……”)为测试案例,我们提出一种跨四种阅读范式(眼动追踪、单向与双向自速阅读、Maze)的人类阅读行为潜变量混合模型。该模型区分了歧义概率、歧义代价与重分析代价,并通过纳入分心阅读试次,得到更真实的加工代价估计。模型能有效再现重读行为、理解问答与语法判断的实证模式。交叉验证显示,该混合模型在预测人类阅读模式及试次结束任务数据方面,优于基于GPT-2突变值的无混合模型。研究对后续工作具有启示意义。
原文摘要 · Abstract (English)
Using temporarily ambiguous garden-path sentences ("While the team trained the striker wondered ...") as a test case, we present a latent-process mixture model of human reading behavior across four different reading paradigms (eye tracking, uni- and bidirectional self-paced reading, Maze). The model distinguishes between garden-path probability, garden-path cost, and reanalysis cost, and yields more realistic processing cost estimates by taking into account trials with inattentive reading. We show that the model is able to reproduce empirical patterns with regard to rereading behavior, comprehension question responses, and grammaticality judgments. Cross-validation reveals that the mixture model also has better predictive fit to human reading patterns and end-of-trial task data than a mixture-free model based on GPT-2-derived surprisal values. We discuss implications for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。