arXiv:2604.10724cs.CL2026-04ACL被引 3

发现重要角色在文本中出现时更意外,且能提升周围内容的可预测性。

Expect the Unexpected? Testing the Surprisal of Salient Entities

  • 通过70,000个标注实体和新方法,研究重要性对信息意外度的影响。
  • 重要实体本身意外度更高,同时让周围内容更易预测。
  • 该现象在主题一致文本中最强,对话类文本中最弱,适合语言认知研究者。

以往研究显示,信息意外度(surprisal)在文档整体上分布较均匀,但受句法与语篇结构影响会产生局部偏差。然而,现有工作普遍忽略话语参与者的重要性。本文利用70,000个跨16种语体的英语文本人工标注提及,结合新颖的最小对提示法,探究话语中实体全局显著性与意外度的关系。结果表明,即使控制位置、长度和嵌套等混淆因素,高显著性实体的意外度仍显著高于非显著实体;且当显著实体作为提示时,会系统性降低其周围内容的意外度,提升文档层面的可预测性。该效应随语体变化:在主题连贯文本中最强,在对话类文本中最弱。研究将统一信息密度(UID)竞争压力框架拓展,揭示全局实体显著性是塑造语篇信息分布的关键机制。

原文摘要 · Abstract (English)

Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across documents overall, local discrepancies can arise due to functional pressures corresponding to syntactic and discourse structural constraints. However, work thus far has largely disregarded the relative salience of discourse participants. We fill this gap by studying how overall salience of entities in discourse relates to surprisal using 70K manually annotated mentions across 16 genres of English and a novel minimal-pair prompting method. Our results show that globally salient entities exhibit significantly higher surprisal than non-salient ones, even controlling for position, length, and nesting confounds. Moreover, salient entities systematically reduce surprisal for surrounding content when used as prompts, enhancing document-level predictability. This effect varies by genre, appearing strongest in topic-coherent texts and weakest in conversational contexts. Our findings refine the UID competing pressures framework by identifying global entity salience as a mechanism shaping information distribution in discourse.

语言模型信息密度语篇分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。