arXiv:2605.21369cs.CL2026-05中稿 · CODI-CRAC 2026被引 4

聚焦跨句长距离指代,扩展多语言数据集提升指代消解性能

Findings of the Fifth Shared Task on Multilingual Coreference Resolution: Expanding Datasets for Long-Range Entities

  • 引入长距离指代链定义,强化跨句子指代识别能力
  • 新增5个数据集、2种语言,覆盖19语种共27个数据集
  • 大模型首次参与即展现潜力,预示未来挑战传统方法

本文介绍与CODI-CRAC 2026研讨会联合举办的第五届多语言指代消解共享任务。该任务要求参赛者构建能够识别提及并进行基于身份的指代聚类的系统。2026年版特别强调长距离实体——即跨越大量词语和句子的指代链。任务在语言覆盖范围上进一步拓展,新增五个数据集及两种语言,均基于CorefUD 1.4版本,该版本为27个数据集组成的统一多语言语料库,涵盖19种语言。共有十套系统参与,其中包括四种基于大模型的方法(三种微调模型和一种少样本方法)。尽管传统系统仍保持领先,但大模型展现出显著潜力,预示其在未来版本中可能挑战现有主流方法。

原文摘要 · Abstract (English)

This paper describes the fifth edition of the Shared Task on Multilingual Coreference Resolution, held in conjunction with the CODI-CRAC 2026 workshop. Building on previous iterations, the task required participants to develop systems capable of mention identification and identity-based coreference clustering. The 2026 edition specifically emphasizes long-range entities, defined as coreferential chains spanning significant distances, across many words and sentences. The task expanded its linguistic scope by incorporating five new datasets and two additional languages. These additions leverage version 1.4 of CorefUD, a harmonized multilingual collection comprising 27 datasets in 19 languages. In total, ten systems participated, including four LLM-based approaches (three fine-tuned models and one few-shot approach). While traditional systems still maintained their lead, LLMs demonstrated significant potential, suggesting they may soon challenge established approaches in future editions.

指代消解多语言长距离指代大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。