构建3部法语长篇小说共28.5万词的指代消解数据集
The Elephant in the Coreference Room: Resolving Coreference in Full-Length French Fiction Works
- 构建覆盖三部长篇小说的细粒度指代标注数据集
- 模型在超28万词长文本上表现良好,支持长链指代解析
- 可辅助分析虚构角色性别,适用于文学与NLP双重场景
尽管指代消解在计算文学研究中日益受到关注,但完整标注的长篇文档数据集仍然极为稀缺。本文引入一个包含三部完整法语小说的新标注语料库,总词数超过285,000词。与以往聚焦短文本的数据集不同,该语料库针对长篇复杂文学作品的挑战,支持在长引用链背景下评估指代消解模型。我们提出一种模块化指代消解流水线,支持细粒度错误分析。结果表明,该方法具有竞争力且能有效扩展至长文档。最后,我们展示了其在推断虚构角色性别方面的应用价值,凸显其对文学分析及下游自然语言处理任务的重要性。
原文摘要 · Abstract (English)
While coreference resolution is attracting more interest than ever from computational literature researchers, representative datasets of fully annotated long documents remain surprisingly scarce. In this paper, we introduce a new annotated corpus of three full-length French novels, totaling over 285,000 tokens. Unlike previous datasets focused on shorter texts, our corpus addresses the challenges posed by long, complex literary works, enabling evaluation of coreference models in the context of long reference chains. We present a modular coreference resolution pipeline that allows for fine-grained error analysis. We show that our approach is competitive and scales effectively to long documents. Finally, we demonstrate its usefulness to infer the gender of fictional characters, showcasing its relevance for both literary analysis and downstream NLP tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。