arXiv:2510.13494cs.CLcs.AI2025-10EMNLP被引 8

清理叙事问答数据集噪声,用高质量文学文本评测大模型理解力

LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA

  • 构建基于文学作品的高质量问答子集,人工与大模型双重验证
  • 发现传统评分指标相关性低,小规模大模型打分更接近人类判断
  • 适合评估长文本理解能力的模型,尤其关注文学叙事理解

在叙事文本上进行问答(QA)对当前系统构成独特挑战,需要深入理解长而复杂的文档。然而,该领域最广泛使用的基准NarrativeQA因文档噪声和问题-答案对缺陷而可靠性受限。本文提出LiteraryQA,一个专注于文学作品的高质量NarrativeQA子集。通过人机协同验证流程,我们识别并修正低质量的问答样本,同时去除源文档中的冗余内容。进一步开展自动评估指标的元评估,揭示所有n-gram类指标与人类判断的系统级相关性较低;而即使使用小型开源大模型作为评判者,其评分也能与人类排名高度一致。最后,我们在LiteraryQA上对一系列长上下文大模型进行了基准测试。代码与数据已公开于https://github.com/SapienzaNLP/LiteraryQA。

原文摘要 · Abstract (English)

Question Answering (QA) on narrative text poses a unique challenge to current systems, requiring a deep understanding of long, complex documents. However, the reliability of NarrativeQA, the most widely used benchmark in this domain, is hindered by noisy documents and flawed QA pairs. In this work, we introduce LiteraryQA, a high-quality subset of NarrativeQA focused on literary works. Using a human- and LLM-validated pipeline, we identify and correct low-quality QA samples while removing extraneous text from source documents. We then carry out a meta-evaluation of automatic metrics to clarify how systems should be evaluated on LiteraryQA. This analysis reveals that all n-gram-based metrics have a low system-level correlation to human judgment, while LLM-as-a-Judge evaluations, even with small open-weight models, can strongly agree with the ranking identified by humans. Finally, we benchmark a set of long-context LLMs on LiteraryQA. We release our code and data at https://github.com/SapienzaNLP/LiteraryQA.

叙事理解文本评测大模型评估文学问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。