arXiv:2602.05758cs.CL2026-02被引 10

用密集奖励强化长文本推理,让大模型更会思考。

LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards

  • 引入动态'想-读'机制,边思考边查文档
  • 在LongBench v2上提升9%,多任务表现稳定
  • 适合需要深度分析长文档的科研与工程场景

强化学习已成为大模型推理的关键驱动力。在长上下文场景(如长对话理解、结构化数据解析)中,挑战不仅在于处理大量文本,更在于进行严谨推演。现有方法多依赖数据合成或架构改进,但仅用稀疏的最终结果奖励效果有限,因粗粒度信号难以有效引导复杂的长上下文推理。为此,我们提出LongR,一个统一框架,通过集成动态'想-读'机制(交替进行推理与文档查阅)和基于相对信息增益的上下文密度奖励,量化相关文档的实用价值。实验表明,LongR在LongBench v2上实现9%性能提升,并在RULER和InfiniteBench上持续优化,对DAPO、GSPO等多种强化学习算法均有稳定增益。进一步分析揭示了推理链长度对效率的影响及模型抗干扰能力。

原文摘要 · Abstract (English)

Reinforcement Learning has emerged as a key driver for LLM reasoning. This capability is equally pivotal in long-context scenarios--such as long-dialogue understanding and structured data analysis, where the challenge extends beyond consuming tokens to performing rigorous deduction. While existing efforts focus on data synthesis or architectural changes, recent work points out that relying solely on sparse, outcome-only rewards yields limited gains, as such coarse signals are often insufficient to effectively guide the complex long-context reasoning. To address this, we propose LongR, a unified framework that enhances long-context performance by integrating a dynamic "Think-and-Read" mechanism, which interleaves reasoning with document consultation, with a contextual density reward based on relative information gain to quantify the utility of the relevant documents. Empirically, LongR achieves a 9% gain on LongBench v2 and consistent improvements on RULER and InfiniteBench, demonstrating robust efficiency in navigating extensive contexts. Furthermore, LongR consistently enhances performance across diverse RL algorithms (e.g., DAPO, GSPO). Finally, we conduct in-depth analyses to investigate the impact of reasoning chain length on efficiency and the model's robustness against distractors.

长文本推理强化学习大模型信息增益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。