arXiv:2607.21574cs.CL2026-07

surprisal理论本质是同义反复,无法验证。

Surprisal Theory is Tautological (without Rational Grounding)

  • 任何心理难度数据都能找到对应语言模型使其呈线性关系
  • 现有模型无法解释人类处理难度,因无额外约束条件
  • 需引入理性机制如记忆或目标约束来打破循环

惊奇度理论认为语言单位在语境中的认知难度是其在某种语言模型下的惊奇度的仿射函数。本文论证该主张本质上是同义反复:在温和技术条件下,对任意非负难度度量,总存在一个语言模型使其惊奇度与难度呈仿射关系。因此,若无对语言模型的额外约束,该理论无法做出可证伪的预测。这一同义反复长期被掩盖,源于过去二十年心理语言学隐含假设——相关语言模型即训练语料分布,提升语料拟合度即提升行为预测能力。但近期实证研究已动摇此假设,证明更优语料模型反而可能降低对处理难度的预测性能。结论是,必须引入理性干预,即语言模型应基于非经验性的理解者模型(如记忆限制或处理目标),而非依赖待解释的行为数据。

原文摘要 · Abstract (English)

Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.

认知科学语言模型理论批判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。