arXiv:2505.02142cs.CL2025-05被引 1

用低成本离线强化学习提升大模型推理能力

Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study

  • 采用离线强化学习中的DPO及改进版LD-DPO方法
  • 在多个基准上平均提升3.3%,难题集提升10.1%
  • 适合追求高效推理优化的研究者和开发者

尽管大语言模型(LLMs)在长上下文推理方面取得显著进展,主要依赖在线强化学习(Online RL)方法,但这类方法计算成本高、复杂度大。相比之下,更简单且经济的离线强化学习(Offline RL)方法尚未被充分探索。为填补这一空白,本文研究了离线强化学习方法——直接偏好优化(DPO)及其长度无关变体LD-DPO——在提升LLM推理能力方面的有效性。在多个推理基准上的广泛实验表明,这些更简单的离线方法显著提升了模型性能,平均提升达3.3%,尤其在具有挑战性的Arena-Hard基准上提升高达10.1%。此外,我们分析了DPO对输出长度的敏感性,强调推理长度增长应与语义丰富性匹配,盲目拉长可能损害性能。本文详述了数据处理与训练方法,为开发更低成本的离线强化学习策略提供了实证依据与实践洞见。

原文摘要 · Abstract (English)

Despite significant advances in long-context reasoning by large language models (LLMs), primarily through Online Reinforcement Learning (RL) methods, these approaches incur substantial computational costs and complexity. In contrast, simpler and more economical Offline RL methods remain underexplored. To address this gap, we investigate the effectiveness of Offline RL methods, specifically Direct Preference Optimization (DPO) and its length-desensitized variant LD-DPO, in enhancing the reasoning capabilities of LLMs. Extensive experiments across multiple reasoning benchmarks demonstrate that these simpler Offline RL methods substantially improve model performance, achieving an average enhancement of 3.3\%, with a particularly notable increase of 10.1\% on the challenging Arena-Hard benchmark. Furthermore, we analyze DPO's sensitivity to output length, emphasizing that increasing reasoning length should align with semantic richness, as indiscriminate lengthening may adversely affect model performance. We provide comprehensive descriptions of our data processing and training methodologies, offering empirical evidence and practical insights for developing more cost-effective Offline RL approaches.

大模型推理离线强化学习偏好优化效率提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。