arXiv:2512.07134cs.CL2025-12被引 3

构建首个覆盖16种语体的英语指代衔接语料库,助力指代消解研究

GUMBridge: a Corpus for Varieties of Bridging Anaphora

  • 构建涵盖16类语体的英文指代衔接语料库,标注细粒度指代类型
  • 在3项任务上测试主流大模型,发现指代消解与类型分类仍具挑战
  • 为指代研究提供丰富数据支持,适合自然语言理解方向研究者

指代衔接是一种语篇现象,其中某实体的指代依赖于前文非同一实体进行理解,例如‘有一栋房子。门是红色的’,此时‘门’特指前述房子的门。尽管已有多个英文指代衔接资源,但多数规模小、覆盖有限且语体单一。本文介绍GUMBridge,一个包含16种多样语体的英文指代衔接新语料库,既涵盖广泛现象,又提供细粒度的指代类型标注。我们评估了标注质量,并报告了开放与闭源主流大语言模型在三项基础任务上的表现,结果显示即使在大模型时代,指代消解与类型分类仍是困难的自然语言处理任务。

原文摘要 · Abstract (English)

Bridging is an anaphoric phenomenon where the referent of an entity in a discourse is dependent on a previous, non-identical entity for interpretation, such as in "There is 'a house'. 'The door' is red," where the door is specifically understood to be the door of the aforementioned house. While there are several existing resources in English for bridging anaphora, most are small, provide limited coverage of the phenomenon, and/or provide limited genre coverage. In this paper, we introduce GUMBridge, a new resource for bridging, which includes 16 diverse genres of English, providing both broad coverage for the phenomenon and granular annotations for the subtype categorization of bridging varieties. We also present an evaluation of annotation quality and report on baseline performance using open and closed source contemporary LLMs on three tasks underlying our data, showing that bridging resolution and subtype classification remain difficult NLP tasks in the age of LLMs.

指代消解语料库自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。