arXiv:2507.22209cs.CL2025-07被引 2

用更精确方法重估词汇熵,发现常用估算方式会误导心理语言学研究。

How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?

  • 通过蒙特卡洛模拟计算跨分词的词汇真实熵值
  • 实验证明首词元熵与真实熵在阅读时间上表现差异显著
  • 提醒研究者慎用简化估算,尤其关注模型生成文本时

上下文熵是衡量词汇预期处理难度的心理语言学指标。近期研究尝试将其作为突兀度效应的补充。为方便起见,通常基于语言模型对首个子词元的概率分布来估算熵值。然而这种近似会导致熵值低估并引入偏差。为此,本文采用蒙特卡洛(MC)方法生成可跨越可变数量分词的真实词汇熵估计。在阅读时间上的回归实验显示,首词元熵与MC熵结果存在明显分歧,提示在使用首词元近似时需谨慎。

原文摘要 · Abstract (English)

Contextual entropy is a psycholinguistic measure capturing the anticipated difficulty of processing a word just before it is encountered. Recent studies have tested for entropy-related effects as a potential complement to well-known effects from surprisal. For convenience, entropy is typically estimated based on a language model's probability distribution over a word's first subword token. However, this approximation results in underestimation and potential distortion of true word entropy. To address this, we generate Monte Carlo (MC) estimates of word entropy that allow words to span a variable number of tokens. Regression experiments on reading times show divergent results between first-token and MC word entropy, suggesting a need for caution in using first-token approximations of contextual entropy.

心理语言学熵估计自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。