arXiv:2605.06506cs.CL2026-05

语言模型预测难易度与隐喻新颖性相关性,实则受词频干扰。

The Frequency Confound in Language-Model Surprisal and Metaphor Novelty

  • 用不同词频指标分析隐喻新颖性评分,发现词频影响更大
  • 模型训练早期惊喜度与新颖性相关最强,后期下降
  • 提示当前最佳设置可能误把词频当预测难易度

语言模型(LM)的惊喜度广泛用作上下文可预测性的代理指标,并被报告与隐喻新颖性判断相关。然而,惊喜度与词汇频率紧密交织。本文通过两种不同的词频度量,研究其在隐喻新颖性评分中的交互作用。分析了八个Pythia模型规模和154个训练检查点的惊喜度估计。在各种设置下,词频对隐喻新颖性的预测能力均强于惊喜度。在训练阶段中,惊喜度与新颖性的关联在早期达到峰值,随后下降,这与惊喜度与频率关联的同步上升趋势一致。这些结果表明,常报告的最优语言模型惊喜度设置可能错误地将上下文可预测性与隐喻新颖性及处理难度关联,而词汇频率可能是主要潜在因素。

原文摘要 · Abstract (English)

Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgments. However, surprisal is tightly intertwined with lexical frequency. We explore this interaction on metaphor novelty ratings using two different word frequency measures. We analyse surprisal estimates from eight Pythia model sizes and 154 training checkpoints. Across settings, word frequency is a stronger predictor of metaphor novelty than surprisal. Across training stages, the surprisal--novelty association peaks at an early stage and then falls again, mirroring a similarly timed increase in the surprisal--frequency association. These results suggest that the often-reported optimal LM surprisal settings may incorrectly associate contextual predictability with metaphor novelty and processing difficulty, whereas lexical frequency may be the major underlying factor.

语言模型隐喻词频心理学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。