arXiv:2601.09886cs.CL2026-01ACL被引 1

语言模型预测比填空实验更准,因它能区分细微语义差异。

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

  • 用语言模型概率替代人工填空数据做预测
  • 模型对低频词和近义词的区分更精准
  • 适合研究人类语言理解中预测机制的细节

一个词的可预测性可通过两种方式衡量:基于人类填空任务的响应或使用语言模型(LM)的概率。当作为认知加工负担的预测因子时,语言模型的概率优于填空数据得出的概率。但需确认其优势源于合理原因,因为不同预测器可能导致关于预测在语言理解中作用的不同结论。本文提供证据支持三个假设:语言模型概率不因分辨率低而失效,能区分语义相近词汇,并准确为低频词赋值。这些结果呼吁提升填空研究的分辨率,并检验人类预测是否同样敏感于模型所捕捉的细微差异。

原文摘要 · Abstract (English)

How predictable a word is can be quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs).When used as predictors of processing effort, LM probabilities outperform probabilities derived from cloze data. However, it is important to establish that LM probabilities do so for the right reasons, since different predictors can lead to different scientific conclusions about the role of prediction in language comprehension. We present evidence for three hypotheses about the advantage of LM probabilities: not suffering from low resolution, distinguishing semantically similar words, and accurately assigning probabilities to low-frequency words. These results call for efforts to improve the resolution of cloze studies, coupled with experiments on whether human-like prediction is also as sensitive to the fine-grained distinctions made by LM probabilities.

语言模型可预测性认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。