arXiv:2602.18217cs.CL2026-02中稿 · CoNLL 2026被引 1

用信息论量化句子理解中的记忆负担,更精准预测阅读时长。

Information-Theoretic Storage Cost in Sentence Comprehension

  • 基于信息论计算前文对后续语境的预测信息量,连续可测。
  • 能复现中心嵌套和定语从句的加工不对称现象。
  • 适合研究语言认知机制或模型可解释性的学者使用。

实时句子理解对工作记忆造成显著负荷,听者需保留上下文以预测后续内容。现有负荷度量多依赖符号语法,对句法预测分配离散、均等的成本。本文提出一种基于信息论的处理存储成本度量方法,即在不确定性下,先前词语携带的关于未来语境的信息量。该方法为连续、概率化、理论中立,且可从预训练神经语言模型中估计。通过三个英语分析验证:(i)恢复中心嵌套与定语从句中的经典加工不对称;(ii)与句法标注语料中的传统语法成本度量相关;(iii)在两个大规模自然语料数据集上,其对阅读时长方差的预测能力超越传统信息基基线模型。代码已开源。

原文摘要 · Abstract (English)

Real-time sentence comprehension imposes a significant load on working memory, as comprehenders must maintain contextual information to anticipate future input. While measures of such load have played an important role in psycholinguistic theories, they have largely been formalized using symbolic grammars, which assign discrete, uniform costs to syntactic predictions. This study proposes a measure of processing storage cost based on an information-theoretic formalization, as the amount of information previous words carry about future context, under uncertainty. Unlike previous discrete, grammar-based metrics, this measure is continuous, probabilistic, theory-neutral, and can be estimated from pre-trained neural language models. The validity of this approach is demonstrated through three analyses in English: our measure (i) recovers well-known processing asymmetries in center embeddings and relative clauses, (ii) correlates with a grammar-based storage cost in a syntactically-annotated corpus, and (iii) predicts reading-time variance in two large-scale naturalistic datasets over and above baseline models with traditional information-based predictors. Our code is available at https://github.com/kohei-kaji/info-storage.

语言认知信息论神经语言模型阅读时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。