纯强化学习可让大模型自发具备过程评估能力,无需依赖外部评分模型。
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
- 通过纯强化学习训练,模型自动生成并评估解题过程。
- 在DeepSeek-R1等模型上,现有PRM表现甚至不如简单投票基准。
- 提出自生成评分框架Self-PRM,但对难题仍存在低精度问题。
大语言模型(LLMs)的推理能力发展是当前研究的关键前沿,强化学习(RL)与过程奖励模型(PRM)被视为主流方法。然而,实证研究表明,专注于数学解题的纯强化学习训练可逐步提升推理能力,无需依赖PRM,挑战了过程监督的必要性。本研究系统考察了强化学习与PRM能力的关系,发现解题能力与过程评估能力在纯强化学习中协同演化。尤其值得注意的是,当前PRM在如DeepSeek-R1和QwQ-32B等先进模型上的表现,甚至低于多数投票等简单基线。为此,我们提出Self-PRM——一种模型自主评估并重排解题方案的内省框架。尽管该方法在基准测试中持续提升准确率(尤其样本量较大时),分析显示其在难题上精度仍低于10%,常将错误解误判为正确。这表明需进一步扩大强化学习规模以优化奖励对齐与内省准确性。总体而言,本研究提示:复杂推理的提升未必依赖PRM,纯强化学习不仅增强解题能力,还自然催生强健的流程评估能力。
原文摘要 · Abstract (English)
The development of reasoning capabilities represents a critical frontier in large language models (LLMs) research, where reinforcement learning (RL) and process reward models (PRMs) have emerged as predominant methodological frameworks. Contrary to conventional wisdom, empirical evidence from DeepSeek-R1 demonstrates that pure RL training focused on mathematical problem-solving can progressively enhance reasoning abilities without PRM integration, challenging the perceived necessity of process supervision. In this study, we conduct a systematic investigation of the relationship between RL training and PRM capabilities. Our findings demonstrate that problem-solving proficiency and process supervision capabilities represent complementary dimensions of reasoning that co-evolve synergistically during pure RL training. In particular, current PRMs underperform simple baselines like majority voting when applied to state-of-the-art models such as DeepSeek-R1 and QwQ-32B. To address this limitation, we propose Self-PRM, an introspective framework in which models autonomously evaluate and rerank their generated solutions through self-reward mechanisms. Although Self-PRM consistently improves the accuracy of the benchmark (particularly with larger sample sizes), analysis exposes persistent challenges: The approach exhibits low precision (<10\%) on difficult problems, frequently misclassifying flawed solutions as valid. These analyses underscore the need for continued RL scaling to improve reward alignment and introspective accuracy. Overall, our findings suggest that PRM may not be essential for enhancing complex reasoning, as pure RL not only improves problem-solving skills but also inherently fosters robust PRM capabilities. We hope these findings provide actionable insights for building more reliable and self-aware complex reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。