arXiv:2606.27598cs.CLcs.AI2026-06

用合成叙事提升长尾实体类型识别准确率

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing

论文配图:Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing
图 1 · 摘自论文原文
  • 构建带叙事上下文的超细粒度实体类型数据集
  • 叙事上下文使长尾类型识别准确率显著提升
  • 可控叙事设计揭示真实文本中隐含的消歧信号

超细粒度实体类型(UFET)为实体提及分配高度具体的类型,但现有方法在长尾类型上表现不佳。我们推测其关键限制在于依赖句级上下文,而消歧证据常分散于多句中。由于现有UFET资源均为句级,验证该假设困难。本文提出Narrative-UFET,通过自动生成连贯短叙事来扩展UFET,使每个实体提及对应一段叙事。设计两种变体:保持类型不变(Maintain)与类型变化(Change)。实验表明,叙事上下文在长尾类型上持续优于句级基线,其中Change变体提供更强信号。与自然语境对比显示,合成叙事带来更大提升,说明可控叙事构造可揭示真实文本中隐含的语义线索。仍存在较大改进空间,提示了话语建模与叙事生成的新方向。

原文摘要 · Abstract (English)

Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level context, since disambiguating evidence is often spread across multiple sentences. Testing this has been difficult because all existing UFET resources are sentence-level. We present Narrative-UFET, a controlled extension of UFET in which each entity mention is paired with an automatically generated short, coherent narrative. Synthesizing narratives lets us isolate the effect of specific discourse properties. We experiment with two paired variants: one in which the entity's type is held constant across the narrative (Maintain) and one in which it shifts (Change). We show that narrative context yields consistent improvements on long-tail types over sentence-level baselines, with the Change variant providing the stronger signal. A comparison against naturally occurring contexts shows that synthetic narratives yield stronger gains, indicating that controlled discourse construction can surface signals that real text leaves implicit. Substantial room for improvement remains, suggesting open directions in both discourse modeling and narrative construction.

实体类型叙事生成长尾识别语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。