构建首个百万级字词书籍级指代消解数据集,推动长文本指代研究
BOOKCOREF: Coreference Resolution at Book Scale
- 自动化流水线生成高质量书籍级指代标注
- 平均文档超20万字词,模型在书中测试提升20点F1
- 揭示现有模型在长文本中性能显著下降的挑战
指代消解系统通常在中小规模文档上评估。然而,现有基准如LitBank在长度上仍有限,无法充分检验系统在书籍尺度(即指代跨度达数十万词)下的能力。为此,我们提出一种新型自动化流水线,可对完整叙事文本进行高质量指代标注,并基于此构建首个书籍级指代消解基准BOOKCOREF,其平均文档长度超过20万词。实验表明该流程稳健,所生成资源使当前长文档指代系统在全书评估中最高提升20点CoNLL-F1。此外,我们揭示了前所未有的书籍尺度带来的新挑战,指出当前模型在长文本中的表现远低于小文档。数据与代码已开源,以促进书籍级指代消解系统的发展。
原文摘要 · Abstract (English)
Coreference Resolution systems are typically evaluated on benchmarks containing small- to medium-scale documents. When it comes to evaluating long texts, however, existing benchmarks, such as LitBank, remain limited in length and do not adequately assess system capabilities at the book scale, i.e., when co-referring mentions span hundreds of thousands of tokens. To fill this gap, we first put forward a novel automatic pipeline that produces high-quality Coreference Resolution annotations on full narrative texts. Then, we adopt this pipeline to create the first book-scale coreference benchmark, BOOKCOREF, with an average document length of more than 200,000 tokens. We carry out a series of experiments showing the robustness of our automatic procedure and demonstrating the value of our resource, which enables current long-document coreference systems to gain up to +20 CoNLL-F1 points when evaluated on full books. Moreover, we report on the new challenges introduced by this unprecedented book-scale setting, highlighting that current models fail to deliver the same performance they achieve on smaller documents. We release our data and code to encourage research and development of new book-scale Coreference Resolution systems at https://github.com/sapienzanlp/bookcoref.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。