提出R-Horizon评估与提升大模型长程推理能力的新方法
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
- 通过查询构造激发模型长程思维链行为
- 发现顶级模型在长程任务中性能显著下降,有效推理长度有限
- 适用于需要多步复杂推理的研究者和开发者
近期测试时扩展的推理模型(如OpenAI o1、DeepSeek-R1)通过长思维链(CoT)取得了显著进步。然而现有基准主要关注即时单阶段任务,难以评估模型在复杂长程场景中的表现。为弥补这一不足,我们提出R-HORIZON,通过查询构造激发大推理模型(LRM)的长程推理行为。基于此构建了一个长程推理基准,包含跨多步、相互依赖的复杂任务。全面评估显示,即使最先进的LRM也出现明显性能退化。分析表明,模型有效推理长度有限,且难以合理分配思考资源。针对此问题,利用R-HORIZON生成长程推理数据,结合验证奖励强化学习(RLVR)。相比单阶段训练,该方法不仅显著提升多阶段推理表现,还在标准任务上实现AIME2024得分提升7.5点。R-HORIZON成为一种可扩展、可控且低成本的增强与评估长程推理能力的范式。
原文摘要 · Abstract (English)
Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek-R1) have led to remarkable improvements through long Chain-of-Thought (CoT). However, existing benchmarks mainly focus on immediate, single-horizon tasks, failing to adequately evaluate models' ability to understand and respond to complex, long-horizon scenarios. To address this incomplete evaluation of Large Reasoning Models (LRMs), we propose R-HORIZON, a method designed to stimulate long-horizon reasoning behaviors in LRMs through query composition. Based on R-HORIZON, we construct a long-horizon reasoning benchmark, comprising complex multi-step reasoning tasks with interdependent problems that span long reasoning horizons. Through comprehensive evaluation of LRMs using the R-HORIZON benchmark, we find that even the most advanced LRMs suffer significant performance degradation. Our analysis reveals that LRMs exhibit limited effective reasoning length and struggle to allocate thinking budget across multiple problems appropriately. Recognizing these limitations, we use R-HORIZON to construct long-horizon reasoning data for reinforcement learning with verified rewards (RLVR). Compared to training with single-horizon data, RLVR with R-HORIZON not only substantially improves performance on the multi-horizon reasoning tasks, but also promotes accuracy on standard reasoning tasks, with an increase of 7.5 on AIME2024. These results position R-HORIZON as a scalable, controllable, and low-cost paradigm for enhancing and evaluating the long-horizon reasoning capabilities of LRMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。