arXiv:2410.10372cs.CLcs.AI2024-10EMNLP被引 6

构建书本角色分析数据集,推动长文本叙事理解研究

BookWorm: A Dataset for Character Description and Analysis

  • 构建包含人类撰写的角色描述与分析的书本级数据集
  • 检索式方法在角色描述与分析任务中优于分层处理方法
  • 基于共指消解的检索微调模型生成最准确的角色描述

角色是所有故事的核心,推动情节并吸引读者。本文研究长篇书籍中角色的理解,这类作品包含复杂叙事和众多交互角色。定义两个任务:角色描述(生成简明事实性简介)与角色分析(深入解读人物发展、性格与社会背景)。提出BookWorm数据集,将古腾堡计划中的书籍与人工撰写的角色描述及分析配对。利用该数据集,在零样本与微调设置下评估先进长上下文模型,采用基于检索与分层处理两种方式处理书本级输入。结果表明,检索式方法在两项任务中均优于分层方法;且基于共指消解的检索微调模型在事实性与蕴含性指标上表现最佳。我们希望本数据集、实验与分析能激发更多关于角色驱动叙事理解的研究。

原文摘要 · Abstract (English)

Characters are at the heart of every story, driving the plot and engaging readers. In this study, we explore the understanding of characters in full-length books, which contain complex narratives and numerous interacting characters. We define two tasks: character description, which generates a brief factual profile, and character analysis, which offers an in-depth interpretation, including character development, personality, and social context. We introduce the BookWorm dataset, pairing books from the Gutenberg Project with human-written descriptions and analyses. Using this dataset, we evaluate state-of-the-art long-context models in zero-shot and fine-tuning settings, utilizing both retrieval-based and hierarchical processing for book-length inputs. Our findings show that retrieval-based approaches outperform hierarchical ones in both tasks. Additionally, fine-tuned models using coreference-based retrieval produce the most factual descriptions, as measured by fact- and entailment-based metrics. We hope our dataset, experiments, and analysis will inspire further research in character-based narrative understanding.

角色分析长文本理解数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。