arXiv:2503.03746cs.CLcs.AI2025-03ACL被引 31

让大模型自己评自己,用分步推理提升数学解题能力。

Process-based Self-Rewarding Language Models

  • 引入分步推理与自评机制,让模型逐步优化自身输出。
  • 在多个数学推理数据集上显著提升性能,超越原有自奖励方法。
  • 适合对大模型自我迭代训练感兴趣的科研人员与工程师。

大型语言模型在多种下游任务中表现出色,广泛应用于实际场景。当前训练通常依赖人工标注的偏好数据,受限于人类表现上限。为此提出自奖励方法,即模型自主生成并评估自身输出。然而现有自奖励范式在数学推理任务中效果不佳,甚至导致性能下降。本文提出基于过程的自奖励框架,包含长时推理、分步模型作为裁判以及分步偏好优化。通过迭代式过程自奖励,该方法显著提升模型在多个数学推理基准上的表现,展现出自奖励机制实现超越人类能力的模型推理的巨大潜力。

原文摘要 · Abstract (English)

Large Language Models have demonstrated outstanding performance across various downstream tasks and have been widely applied in multiple scenarios. Human-annotated preference data is used for training to further improve LLMs' performance, which is constrained by the upper limit of human performance. Therefore, Self-Rewarding method has been proposed, where LLMs generate training data by rewarding their own outputs. However, the existing self-rewarding paradigm is not effective in mathematical reasoning scenarios and may even lead to a decline in performance. In this work, we propose the Process-based Self-Rewarding pipeline for language models, which introduces long-thought reasoning, step-wise LLM-as-a-Judge, and step-wise preference optimization within the self-rewarding paradigm. Our new paradigm successfully enhances the performance of LLMs on multiple mathematical reasoning benchmarks through iterative Process-based Self-Rewarding, demonstrating the immense potential of self-rewarding to achieve LLM reasoning that may surpass human capabilities.

自奖励数学推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。