通过聚焦推理过程中的关键步骤,提升大模型的多步推理能力。
GPO: Learning from Critical Steps to Improve LLM Reasoning
- 识别推理轨迹中的关键步骤,定位影响成败的决策点。
- 在关键步骤处重置策略并优先学习,显著提升推理准确率。
- 适用于各类优化方法,适合研究大模型推理机制的学者。
大语言模型在多个领域展现出巨大潜力,近期提出的推理型大模型旨在增强其解决复杂问题的思维能力。尽管已有进展,但提升多步推理能力仍是重大挑战。现有优化方法通常将推理路径整体处理,忽视其中的关键步骤。本文提出引导式关键优化(GPO),通过估计优势函数识别推理轨迹中的关键步骤——模型必须谨慎推进以成功解决问题的节点。GPO在该步骤处重置策略,采样新轨迹,并优先学习这些关键路径。这一聚焦策略使模型能更有效地从推理过程中的关键时刻中学习,从而提升性能。理论分析与实验证明,GPO可与多种优化方法结合,在多个挑战性推理基准上持续显著提升现有方法的表现,验证了其在聚焦生成过程关键环节方面的有效性与通用性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in various domains, showing impressive potential on different tasks. Recently, reasoning LLMs have been proposed to improve the \textit{reasoning} or \textit{thinking} capabilities of LLMs to solve complex problems. Despite the promising results of reasoning LLMs, enhancing the multi-step reasoning capabilities of LLMs still remains a significant challenge. While existing optimization methods have advanced the LLM reasoning capabilities, they often treat reasoning trajectories as a whole, without considering the underlying critical steps within the trajectory. In this paper, we introduce \textbf{G}uided \textbf{P}ivotal \textbf{O}ptimization (GPO), a novel fine-tuning strategy that dives into the reasoning process to enable more effective improvements. GPO first identifies the `critical step' within a reasoning trajectory - a point that the model must carefully proceed to succeed at the problem. We locate the critical step by estimating the advantage function. GPO then resets the policy to the critical step, samples the new rollout and prioritizes the learning process on those rollouts. This focus allows the model to learn more effectively from pivotal moments within the reasoning process to improve the reasoning performance. We demonstrate that GPO is a general strategy that can be integrated with various optimization methods to improve reasoning performance. Besides theoretical analysis, our experiments across challenging reasoning benchmarks show that GPO can consistently and significantly enhance the performance of existing optimization methods, showcasing its effectiveness and generalizability in improving LLM reasoning by concentrating on pivotal moments within the generation process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。