用强化学习提升大模型推理能力,实现测试时规模扩展。
T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling
- 用合成思维链数据初始化模型,结合试错与自验证。
- 通过过采样增强采样多样性,实现强化学习有效训练。
- 无需额外验证,增加推理预算即可提升性能,适合数学推理研究者。
大语言模型在复杂推理任务中表现出色,但现有方法多依赖模仿学习,难以实现有效的测试时规模扩展。尽管强化学习(RL)有望实现自我探索,但近期尝试仅带来有限提升。本文提出T1,通过鼓励探索并理解推理规模扩展来提升RL效果。首先,使用融合试错与自验证的合成思维链数据初始化大语言模型。为扩大强化学习训练规模,通过过采样提升采样多样性。实验表明,以开源大模型为基础的T1展现出推理规模扩展特性,在挑战性数学推理基准上表现更优。更重要的是,我们提出一种简单策略:增加推理预算可直接提升T1性能,无需额外验证。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, existing approaches mainly rely on imitation learning and struggle to achieve effective test-time scaling. While reinforcement learning (RL) holds promise for enabling self-exploration, recent attempts yield modest improvements in complex reasoning. In this paper, we present T1 to scale RL by encouraging exploration and understand inference scaling. We first initialize the LLM using synthesized chain-of-thought data that integrates trial-and-error and self-verification. To scale RL training, we promote increased sampling diversity through oversampling. We demonstrate that T1 with open LLMs as its base exhibits inference scaling behavior and achieves superior performance on challenging math reasoning benchmarks. More importantly, we present a simple strategy to examine inference scaling, where increased inference budgets directly lead to T1's better performance without any additional verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。