多语言指代消解新模型,空节点预测提升9.5%性能
CorPipe at CRAC 2026: Empty Nodes and Cross-Lingual Transfer in Multilingual Coreference Resolution

- 单模型联合预测指代提及、链接与空节点
- 在无约束赛道领先其他系统9.5个百分点
- 支持跨语言零样本推理,代码开源可复现
我们介绍 CorPipe 26,这是在 CRAC 2026 多语言指代消解共享任务中的获奖提交。该任务第五届聚焦生成式大模型与专用系统对比,并新增5个数据集和2种语言。CorPipe 26 是 CorPipe 25 的改进版本,引入一种新变体,能在单一模型中联合预测空节点、提及及指代链接。该系统在大模型赛道上超越所有其他提交2.8个百分点,在无约束赛道上领先9.5个百分点。此外,我们进行了多组消融实验,涵盖不同模型规模、空节点预测方法以及跨语言零样本评估。源代码与训练模型已公开于 https://github.com/ufal/crac2026-corpipe。
原文摘要 · Abstract (English)
We introduce CorPipe 26, our winning submission to the CRAC 2026 Shared Task on Multilingual Coreference Resolution. The fifth edition of this shared task focuses mainly on the comparison of generative LLMs and specialized systems; additionally, 5 more datasets and 2 new languages are introduced. CorPipe 26 is an improved version of CorPipe 25, with a new variant predicting empty nodes together with mentions and coreference links in a single model. Our system outperforms all other submissions in the LLM track by 2.8 percent points and all submissions in the unconstrained track by 9.5 percent points. Furthermore, we perform a series of ablation experiments with different model sizes, empty node prediction methods, and cross-lingual zero-shot evaluation. The source code and the trained models are publicly available at https://github.com/ufal/crac2026-corpipe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。