arXiv:2603.00621cs.CL2026-03中稿 · LREC 2026

整合多领域跨文档共指数据集,统一标准提升研究可复现性。

Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification

  • 构建统一格式的uCDCR数据集,融合实体与事件共指标注。
  • ECB+词汇多样性最低,其共指复杂度处于中等水平。
  • 实验证明联合训练可提升模型泛化能力,适合共指研究者使用。

由于数据集格式不一、标注标准差异以及以事件共指为核心定义(ECR)的倾向,跨文档共指消解(CDCR)研究长期碎片化。本文提出uCDCR,一个将多个公开英文CDCR语料库统一为一致格式的综合性数据集,支持标准化评估。uCDCR包含实体与事件共指,修正了已知不一致,并补充缺失属性,促进可复现研究。通过统一指标分析各数据集词汇特性,如提及项的词汇组成、词汇多样性与歧义度,探讨影响性能的标注原则。结果表明,当前最先进基准ECB+词汇多样性最低,其共指复杂度在所有uCDCR数据集中居中。对比文档与提及分布发现,使用全部uCDCR数据训练评估能显著提升模型泛化能力。同一头词基线在事件和实体任务上表现几乎相同,说明两者均具挑战性,不应仅聚焦于事件共指。数据集与代码已开源:https://huggingface.co/datasets/AnZhu/uCDCR,https://github.com/anastasia-zhukova/uCDCR。

原文摘要 · Abstract (English)

Research in CDCR remains fragmented due to heterogeneous dataset formats, varying annotation standards, and the predominance of the CDCR definition as the event coreference resolution (ECR). To address these challenges, we introduce uCDCR, a unified dataset that consolidates diverse publicly available English CDCR corpora across various domains into a consistent format, which we analyze with standardized metrics and evaluation protocols. uCDCR incorporates both entity and event coreference, corrects known inconsistencies, and enriches datasets with missing attributes to facilitate reproducible research. We establish a cohesive framework for fair, interpretable, and cross-dataset analysis in CDCR and compare the datasets on their lexical properties, e.g., lexical composition of the annotated mentions, lexical diversity and ambiguity metrics, discuss the annotation rules and principles that lead to high lexical diversity, and examine how these metrics influence performance on the same-head-lemma baseline. Our dataset analysis shows that ECB+, the state-of-the-art benchmark for CDCR, has one of the lowest lexical diversities, and its CDCR complexity, measured by the same-head-lemma baseline, lies in the middle among all uCDCR datasets. Moreover, comparing document and mention distributions between ECB+ and uCDCR shows that using all uCDCR datasets for model training and evaluation will improve the generalizability of CDCR models. Finally, the almost identical performance on the same-head-lemma baseline, separately applied to events and entities, shows that resolving both types is a complex task and should not be steered toward ECR alone. The uCDCR dataset is available at https://huggingface.co/datasets/AnZhu/uCDCR, and the code for parsing, analyzing, and scoring the dataset is available at https://github.com/anastasia-zhukova/uCDCR.

共指消解数据集统一自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。