arXiv:2411.11681cs.AIcs.LG2024-11被引 8

通过非线性奖励优化推理链,提升大模型逻辑准确性。

PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment

  • 基于推理步骤数与奖励的非线性关系设计新策略
  • 在6个数学推理数据集上超越主流模型表现
  • 适合需要精准逻辑推理的AI系统开发者

过程监督通过在思维链的每一步提供反馈,提升大语言模型在推理任务中的表现。然而,由于缺乏有效的过程监督方法,即使先进的大模型仍易出现逻辑错误和冗余推理。我们指出,过程监督的有效性显著依赖于推理链的准确性和长度,且这两者与整体奖励分数呈非线性关系。受此启发,我们提出新型过程监督范式PSPO*,系统化地规划从奖励模型训练到策略优化的工作流程,并强调非线性奖励的重要性。基于PSPO*,我们开发了PSPO-WRS,其根据推理步数决定奖励分,并采用修正的Weibull分布进行非线性奖励塑造。在六个数学推理数据集上的实验结果表明,PSPO-WRS持续优于当前主流模型。

原文摘要 · Abstract (English)

Process supervision enhances the performance of large language models in reasoning tasks by providing feedback at each step of chain-of-thought reasoning. However, due to the lack of effective process supervision methods, even advanced large language models are prone to logical errors and redundant reasoning. We claim that the effectiveness of process supervision significantly depends on both the accuracy and the length of reasoning chains. Moreover, we identify that these factors exhibit a nonlinear relationship with the overall reward score of the reasoning process. Inspired by these insights, we propose a novel process supervision paradigm, PSPO*, which systematically outlines the workflow from reward model training to policy optimization, and highlights the importance of nonlinear rewards in process supervision. Based on PSPO*, we develop the PSPO-WRS, which considers the number of reasoning steps in determining reward scores and utilizes an adjusted Weibull distribution for nonlinear reward shaping. Experimental results on six mathematical reasoning datasets demonstrate that PSPO-WRS consistently outperforms current mainstream models.

推理对齐策略优化非线性奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。