用思维推理提升奖励模型,解决复杂推理中的评分难题
Libra: Assessing and Improving Reward Model by Learning to Think
- 设计新型思维导向基准测试Libra Bench,覆盖多样数学难题
- 提出学习思考的生成式奖励模型,实现无参考答案评分
- 适合研究大模型推理与强化学习融合的开发者使用
强化学习显著提升了大语言模型的推理能力,但现有奖励模型在复杂推理场景中表现不佳。主流RL训练依赖规则或参考答案生成奖励,存在两大限制:一是需精细标注参考答案,二是要求输出格式受限。这两点严重制约了数据规模扩展和推理性能持续提升。为此,我们提出一个全面评估与改进奖励模型的框架。首先构建了基于多样化高难度数学问题和先进推理模型的思维导向基准(Libra Bench),以弥补现有基准在推理场景中的不足。进一步提出一种通过‘学习思考’方法改进生成式奖励模型的新策略。基于此,开发出Libra-RM系列具备推理能力的生成式奖励模型,在多个基准上达到领先水平。大量下游实验表明,Libra Bench与实际应用高度相关,且Libra-RM可利用未标注数据进一步提升推理模型性能。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has significantly improved the reasoning ability of large language models. However, current reward models underperform in challenging reasoning scenarios and predominant RL training paradigms rely on rule-based or reference-based rewards, which impose two critical limitations: 1) the dependence on finely annotated reference answer to attain rewards; and 2) the requirement for constrained output format. These limitations fundamentally hinder further RL data scaling and sustained enhancement of model reasoning performance. To address these limitations, we propose a comprehensive framework for evaluating and improving the performance of reward models in complex reasoning scenarios. We first present a reasoning-oriented benchmark (Libra Bench), systematically constructed from a diverse collection of challenging mathematical problems and advanced reasoning models, to address the limitations of existing reward model benchmarks in reasoning scenarios. We further introduce a novel approach for improving the generative reward model via learning-to-think methodologies. Based on the proposed approach, we develop Libra-RM series, a collection of generative reward models with reasoning capabilities that achieve state-of-the-art results on various benchmarks. Comprehensive downstream experiments are conducted and the experimental results demonstrate the correlation between our Libra Bench and downstream application, and the potential of Libra-RM to further improve reasoning models with unlabeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。