arXiv:2602.17424cs.CL2026-02中稿 · ACDSA 2026

重构新闻跨文档指代识别,让不同说法指向同一实体

Diverse Word Choices, Same Reference: Annotating Lexically-Rich Cross-Document Coreference

  • 将指代链视为话语单元,兼容同义与近义关系
  • 重新标注新闻数据集,提升词汇多样性捕捉能力
  • 适合研究媒体话语差异与偏见的学者使用

跨文档指代消解(CDCR)旨在识别并链接相关文档中同一实体或事件的提及,实现以话语参与者为单位的信息整合。然而现有数据集多聚焦事件消解,且对指代定义狭窄,难以应对词汇差异显著的新闻报道分析需求。本文提出针对NewsWCL50数据集的修订标注方案,将指代链视为话语元素(DEs)和分析概念单位,支持身份与近似身份关系,例如将“难民车队”、“寻求庇护者”、“非法入境者”等不同表述关联,使模型能捕捉媒体话语中的词汇多样性和框架差异,同时保持对话语元素的细粒度标注。我们使用统一编码手册重标注了NewsWCL50及ECB+子集,并通过词汇多样性指标与同头词干基线进行评估。结果表明,重标注数据集表现介于原始ECB+与NewsWCL50之间,验证了其在新闻领域平衡、语境敏感的指代消解研究中的有效性。

原文摘要 · Abstract (English)

Cross-document coreference resolution (CDCR) identifies and links mentions of the same entities and events across related documents, enabling content analysis that aggregates information at the level of discourse participants. However, existing datasets primarily focus on event resolution and employ a narrow definition of coreference, which limits their effectiveness in analyzing diverse and polarized news coverage where wording varies widely. This paper proposes a revised CDCR annotation scheme of the NewsWCL50 dataset, treating coreference chains as discourse elements (DEs) and conceptual units of analysis. The approach accommodates both identity and near-identity relations, e.g., by linking "the caravan" - "asylum seekers" - "those contemplating illegal entry", allowing models to capture lexical diversity and framing variation in media discourse, while maintaining the fine-grained annotation of DEs. We reannotate the NewsWCL50 and a subset of ECB+ using a unified codebook and evaluate the new datasets through lexical diversity metrics and a same-head-lemma baseline. The results show that the reannotated datasets align closely, falling between the original ECB+ and NewsWCL50, thereby supporting balanced and discourse-aware CDCR research in the news domain.

指代消解新闻分析话语研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。