多语言指代消解共享任务升级,更贴近真实场景。
Findings of the Third Shared Task on Multilingual Coreference Resolution
- 不提供零回指的黄金标注,提升任务真实性和难度
- 覆盖15种语言的21个数据集,包含历史语言
- 6个系统参赛,推动多语言指代研究进展
本文综述了作为CRAC 2024研讨会一部分举办的第三届多语言指代消解共享任务。与前两届类似,参赛者需构建系统以识别提及并根据指代同一性进行聚类。本届任务进一步向真实应用迈进,不再提供零回指的黄金标注,显著增加任务复杂度与现实性。同时,任务扩展至涵盖更多样化的语言,尤其关注历史语言。训练与评估数据源自多语言统一指代资源CorefUD的1.2版本,涵盖21个数据集、15种语言。共有6个系统参与本次共享任务。
原文摘要 · Abstract (English)
The paper presents an overview of the third edition of the shared task on multilingual coreference resolution, held as part of the CRAC 2024 workshop. Similarly to the previous two editions, the participants were challenged to develop systems capable of identifying mentions and clustering them based on identity coreference. This year's edition took another step towards real-world application by not providing participants with gold slots for zero anaphora, increasing the task's complexity and realism. In addition, the shared task was expanded to include a more diverse set of languages, with a particular focus on historical languages. The training and evaluation data were drawn from version 1.2 of the multilingual collection of harmonized coreference resources CorefUD, encompassing 21 datasets across 15 languages. 6 systems competed in this shared task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。