arXiv:2607.20208cs.CL2026-07

批判性指出大模型困惑度不能随意互换,揭示其背后隐藏的表征与算法选择

surprisal is Not a Theory

论文配图:surprisal is Not a Theory
图 1 · 摘自论文原文
  • 指出语言模型困惑度计算受架构与算法显著影响
  • 实证显示不同模型生成的概率差异显著
  • 提醒研究者警惕将大模型概率视为可互换的误区

Surprisal Theory 常被视作一种计算层级解释(Marr, 1982)。本文认为,尽管该理论常被用于支持计算心理语言学中的‘表征无关’研究,但大型语言模型(LLMs)的黑箱化趋势并未免除使用困惑度指标的研究者在表征层面所作的抉择。事实上,对 LLM 困惑度的无批判使用会掩盖不同模型在表征与算法层级上的隐含假设。通过三项分析,我们表明算法选择与模型架构对语言模型概率的计算具有决定性影响。我们建议希望检验困惑度理论的研究者重新审视将大语言模型概率视为可互换的做法。

原文摘要 · Abstract (English)

Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within computational psycholinguistics, the movement toward black box systems embodied by large language models (LLMs) does not exempt modelers using the surprisal metric from the representational decisions required by computational-level characterizations. In fact, we argue that the uncritical use of LLM-surprisal obfuscates the representational and algorithmic-level commitments of different models. In three analyses, we show that the choice of algorithm and model architecture play significant roles in the computation of language model probabilities. We advise that researchers who wish to test Surprisal Theory re-evaluate the practice of treating large language model probabilities as interchangeable

语言模型困惑度表征认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。