arXiv:2412.15204cs.CLcs.AI2024-12被引 360

LongBench v2评估大模型在超长文本中的深度理解与推理能力。

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

  • 构建涵盖6类任务的503道长文本多选题,上下文达8k至200万词
  • 人类专家在15分钟内仅能达53.7%准确率,最优模型仅50.1%
  • 需增强推理与计算资源才能突破人类表现

本文提出LongBench v2,一个用于评估大语言模型在真实世界多任务中处理超长上下文、实现深度理解与推理能力的基准。该基准包含503道挑战性多项选择题,覆盖单文档问答、多文档问答、长上下文学习、长对话历史理解、代码仓库理解与长结构化数据理解六大任务类别,上下文长度从8,000到200万词不等。数据由近100位具有不同专业背景的高学历人士提供,经自动化与人工双重审核以确保质量与难度。在15分钟时间限制下,人类专家平均准确率为53.7%。最佳直接回答模型准确率仅为50.1%,而采用更长推理链的o1-preview模型达到57.7%,首次超越人类基线4个百分点。结果表明,提升推理能力与增加推理时计算开销对攻克长上下文挑战至关重要。项目开源地址:https://longbench2.github.io。

原文摘要 · Abstract (English)

This paper introduces LongBench v2, a benchmark designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 consists of 503 challenging multiple-choice questions, with contexts ranging from 8k to 2M words, across six major task categories: single-document QA, multi-document QA, long in-context learning, long-dialogue history understanding, code repository understanding, and long structured data understanding. To ensure the breadth and the practicality, we collect data from nearly 100 highly educated individuals with diverse professional backgrounds. We employ both automated and manual review processes to maintain high quality and difficulty, resulting in human experts achieving only 53.7% accuracy under a 15-minute time constraint. Our evaluation reveals that the best-performing model, when directly answers the questions, achieves only 50.1% accuracy. In contrast, the o1-preview model, which includes longer reasoning, achieves 57.7%, surpassing the human baseline by 4%. These results highlight the importance of enhanced reasoning ability and scaling inference-time compute to tackle the long-context challenges in LongBench v2. The project is available at https://longbench2.github.io.

长文本理解多任务评估推理能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。