arXiv:2604.28147cs.CL2026-04ACL被引 2

厘清语言单位定义与预测范围的关系,统一可变单位的意外度计算框架。

On the Proper Treatment of Units in Surprisal Theory

  • 区分语言单位定义与预测区域,避免混淆建模选择。
  • 提出统一框架,支持任意单位体系下的意外度计算。
  • 适合关注语言认知建模或文本分析的学者参考。

意外度理论将人类语言处理努力与下一个语言单元的可预测性联系起来,但实证研究常对‘语言单元’概念表述模糊。实验中通常按语言学动机划分单元(如词),而预训练语言模型则基于固定标记词汇表分配概率,该词汇表通常与实际语言单元不一致。由此导致基于意外度的预测模型隐含依赖于人为设定的处理流程,混淆了两个独立的建模决策:分析单位的定义,以及预测评估区域的选择。本文通过解耦这两个选择,提出一个统一框架,用于在任意单位体系下进行意外度推理。我们认为,意外度分析应明确表达这些选择,并将分词视为实现细节而非科学基础。

原文摘要 · Abstract (English)

Surprisal theory links human processing effort to the predictability of an upcoming linguistic unit, but empirical work often leaves the notion of a unit underspecified. In practice, experimental stimuli are segmented into linguistically motivated units (e.g., words), while pretrained language models assign probability mass to a fixed token alphabet that typically does not align with those units. As a result, surprisal-based predictors depend implicitly on ad hoc procedures that conflate two distinct modeling choices: the definition of the unit of analysis and the choice of regions of interest over which predictions are evaluated. In this paper, we disentangle these choices and give a unified framework for reasoning about surprisal over arbitrary unit inventories. We argue that surprisal-based analyses should make these choices explicit and treat tokenization as an implementation detail rather than a scientific primitive.

语言认知意外度建模框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。