arXiv:2508.09848cs.CLcs.AI2025-08被引 7

测试大模型对长文本的全局理解与推理能力,发现当前模型仍远逊于人类。

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

  • 通过判断前传故事是否符合原著逻辑来评估模型全局理解力。
  • 88% 的题目需整合多个段落信息,模型准确率比人类低15%以上。
  • 模型常错解但答对,推理准确率差距超30%,适合研究长文本推理者关注。

我们提出 PRELUDE,一个通过判断角色前传故事是否与原著正典叙事一致,来评估长上下文理解能力的基准。该任务要求更强的全局理解和深度推理——因前传非原著内容,判断其合理性通常需搜索并整合间接相关的信息。实验显示,88% 的实例需要跨多个部分的信息。结果表明,当前最先进的 LLM 在上下文学习、RAG、领域内训练及商用 DeepResearch 服务中,均落后于人类超过15%。进一步的人类研究表明,模型虽常给出正确答案,但推理过程存在错误,导致推理准确率比人类高出30%以上。这些发现凸显了长上下文理解与推理仍有巨大提升空间。

原文摘要 · Abstract (English)

We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.

长文本理解推理评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。