arXiv:2412.11560cs.CL2024-12被引 7

研究NLP任务对文学角色网络构建的影响,发现命名实体识别和共指消解至关重要。

The Role of Natural Language Processing Tasks in Automatic Literary Character Network Construction

  • 用金标准数据逐步添加错误,评估NLP任务对角色网络质量的影响。
  • 命名实体识别性能直接影响角色检测,仅靠提及会遗漏大量角色共现。
  • 传统NLP管道在召回率上优于大模型端到端方法,适合角色网络构建。

从文学文本中自动提取角色网络通常依赖自然语言处理(NLP)级联流水线。尽管该方法广泛应用,但尚无研究探讨底层NLP任务对其性能的影响。本文针对文学数据集开展研究,聚焦命名实体识别(NER)和共指消解在提取共现网络中的作用。通过使用金标准标注,逐步引入均匀分布的错误,观察其对角色网络质量的影响。结果表明,NER性能随小说类型变化显著,并强烈影响角色检测;仅依赖NER识别的提及会遗漏大量角色共现,需共指消解弥补。此外,与两种基于大语言模型(LLM)的方法(包括一个全端到端方法)对比发现,传统NLP流水线在召回率上表现更优。

原文摘要 · Abstract (English)

The automatic extraction of character networks from literary texts is generally carried out using natural language processing (NLP) cascading pipelines. While this approach is widespread, no study exists on the impact of low-level NLP tasks on their performance. In this article, we conduct such a study on a literary dataset, focusing on the role of named entity recognition (NER) and coreference resolution when extracting co-occurrence networks. To highlight the impact of these tasks' performance, we start with gold-standard annotations, progressively add uniformly distributed errors, and observe their impact in terms of character network quality. We demonstrate that NER performance depends on the tested novel and strongly affects character detection. We also show that NER-detected mentions alone miss a lot of character co-occurrences, and that coreference resolution is needed to prevent this. Finally, we present comparison points with 2 methods based on large language models (LLMs), including a fully end-to-end one, and show that these models are outperformed by traditional NLP pipelines in terms of recall.

角色网络NLP应用共指消解文学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。