arXiv:2506.03090cs.CL2025-06ACL被引 2

用长文本模型补全文学作品引文,测试其文学理解能力。

Literary Evidence Retrieval via Long-Context Language Models

  • 以整部文学作品为上下文,让模型补全缺失引文。
  • 闭源模型准确率达62.5%,超过人类专家的50%。
  • 开源模型仅29.1%准确率,凸显推理能力差距。

现代长文本语言模型对文学小说的理解能力如何?我们通过文学证据检索任务探索这一问题,重新利用That等(2022)提出的RELiC数据集,构建了一个基准测试:将一部原著全文(如《了不起的盖茨比》)提供给大模型,并给出带有缺失引文的文学评论。该设定要求模型在全局叙事推理与细读文本之间切换,模拟人类文学分析过程。我们通过大量筛选与人工验证,构建了包含292个高质量样本的子集。实验表明,近期推理模型如Gemini Pro 2.5可超越人类专家表现(准确率62.5%对比50%),而最佳开源模型仅达29.1%准确率,凸显闭源与开源模型在诠释推理上的显著差距。尽管速度快、表面准确,最强模型仍难以捕捉微妙的文学信号并常出现过度生成,表明将大模型应用于文学分析仍面临严峻挑战。我们已发布数据集与评估代码,以推动该方向研究。

原文摘要 · Abstract (English)

How well do modern long-context language models understand literary fiction? We explore this question via the task of literary evidence retrieval, repurposing the RELiC dataset of That et al. (2022) to construct a benchmark where the entire text of a primary source (e.g., The Great Gatsby) is provided to an LLM alongside literary criticism with a missing quotation from that work. This setting, in which the model must generate the missing quotation, mirrors the human process of literary analysis by requiring models to perform both global narrative reasoning and close textual examination. We curate a high-quality subset of 292 examples through extensive filtering and human verification. Our experiments show that recent reasoning models, such as Gemini Pro 2.5 can exceed human expert performance (62.5% vs. 50% accuracy). In contrast, the best open-weight model achieves only 29.1% accuracy, highlighting a wide gap in interpretive reasoning between open and closed-weight models. Despite their speed and apparent accuracy, even the strongest models struggle with nuanced literary signals and overgeneration, signaling open challenges for applying LLMs to literary analysis. We release our dataset and evaluation code to encourage future work in this direction.

文学理解长文本证据检索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。