arXiv:2505.17315cs.AIcs.CL2025-05NeurIPS被引 7

提升长上下文能力可显著增强模型推理性能,且效果不局限于长输入任务。

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning

  • 在相同训练数据下,增强长上下文能力后模型推理准确率更高
  • 长上下文训练带来的性能提升在短输入任务中依然存在
  • 建议将长上下文能力作为未来模型设计的核心目标

近期语言模型展现出强大的推理能力,但长上下文能力对推理的影响仍缺乏深入研究。本文假设当前推理瓶颈部分源于长上下文能力不足,依据实证观察:(1)更长的上下文窗口通常带来更强的推理表现;(2)推理失败案例与长上下文处理失败模式相似。为验证该假设,我们对比了架构和微调数据相同但长上下文能力不同的模型。结果表明,具备更强长上下文能力的模型在经过监督微调(SFT)后,在多个推理基准上均取得显著更高的准确率。值得注意的是,这种提升在输入长度较短的任务中依然存在,说明长上下文训练具有通用性优势。研究揭示,长上下文建模不仅是处理长输入的需要,更是推理能力的基础支撑。因此,我们主张将长上下文能力作为未来语言模型设计的首要目标。

原文摘要 · Abstract (English)

Recent language models exhibit strong reasoning capabilities, yet the influence of long-context capacity on reasoning remains underexplored. In this work, we hypothesize that current limitations in reasoning stem, in part, from insufficient long-context capacity, motivated by empirical observations such as (1) higher context window length often leads to stronger reasoning performance, and (2) failed reasoning cases resemble failed long-context cases. To test this hypothesis, we examine whether enhancing a model's long-context ability before Supervised Fine-Tuning (SFT) leads to improved reasoning performance. Specifically, we compared models with identical architectures and fine-tuning data but varying levels of long-context capacity. Our results reveal a consistent trend: models with stronger long-context capacity achieve significantly higher accuracy on reasoning benchmarks after SFT. Notably, these gains persist even on tasks with short input lengths, indicating that long-context training offers generalizable benefits for reasoning performance. These findings suggest that long-context modeling is not just essential for processing lengthy inputs, but also serves as a critical foundation for reasoning. We advocate for treating long-context capacity as a first-class objective in the design of future language models.

推理能力长上下文语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。