arXiv:2501.15089cs.CL2025-01被引 46

构建合成长上下文推理数据集,评估大模型在复杂推理中的表现

LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion

  • 通过扩展短文本推理题生成长上下文题目
  • 21个模型在长上下文中性能普遍下降,顶尖模型仍有提升空间
  • 适合研究长上下文推理能力的学者和开发者使用

大型语言模型在理解长上下文输入方面已取得显著进展,但评估其长上下文推理能力的基准测试却滞后。现有基准多集中于狭窄任务或简单推理。为弥补这一差距,我们提出新合成基准LongReason,通过上下文扩展从多样化短上下文推理题生成长上下文问题。LongReason包含794道多选题,涵盖阅读理解、逻辑推理和数学应用题三类任务,具有多样化的推理模式。我们在LongReason上评估了21个LLM,发现多数模型在上下文长度增加时性能显著下降。进一步分析表明,即使顶尖模型在跨任务鲁棒推理方面仍有巨大提升空间。LongReason已开源,地址为https://huggingface.co/datasets/lz1bytedance/LongReason,以支持对大模型长上下文推理能力的全面评估。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable progress in understanding long-context inputs. However, benchmarks for evaluating the long-context reasoning abilities of LLMs fall behind the pace. Existing benchmarks often focus on a narrow range of tasks or those that do not demand complex reasoning. To address this gap and enable a more comprehensive evaluation of the long-context reasoning capabilities of current LLMs, we propose a new synthetic benchmark, LongReason, which is constructed by synthesizing long-context reasoning questions from a varied set of short-context reasoning questions through context expansion. LongReason consists of 794 multiple-choice reasoning questions with diverse reasoning patterns across three task categories: reading comprehension, logical inference, and mathematical word problems. We evaluate 21 LLMs on LongReason, revealing that most models experience significant performance drops as context length increases. Our further analysis shows that even state-of-the-art LLMs still have significant room for improvement in providing robust reasoning across different tasks. We have open-sourced LongReason under https://huggingface.co/datasets/lz1bytedance/LongReason to support the comprehensive evaluation of LLMs' long-context reasoning capabilities.

长上下文推理评估合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。