arXiv:2412.17397cs.LGcs.CV2024-12中稿 · AAAI被引 2

通过自纠正机制提升大模型推理能力,显著改善数学题准确率。

Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning

  • 利用自生成数据进行两阶段训练,增强模型自我修正能力。
  • 在MATH和GSM8K数据集上准确率分别提升4.18%和2.28%。
  • 适合需要高精度数学推理的场景,如自动解题系统。

当前主流方法通过受AlphaZero启发的迭代偏好学习来提升大语言模型(LLMs)的推理能力。本文提出进一步通过内在自纠正机制增强逐步推理能力。我们采用分两阶段的训练流程:第一阶段,仅依赖模型自身预测生成的数据,在内在自纠正机制下增强其自我验证能力;第二阶段,基于第一阶段获得的优化自纠正策略,应用基线逐步偏好学习。在算术推理任务评估中,该方法在MATH数据集上准确率提升至71.34%(+4.18%),优于OpenMath2-Llama3.1-8B与dart-math-mistral-7b-uniform;在GSM8K数据集上准确率达86.76%(+2.00%),优于Llama-3.1-8B-Instruct与Mistral-7B-Instruct-v0.1。

原文摘要 · Abstract (English)

With current state-of-the-art approaches aimed at enhancing the reasoning capabilities of Large Language Models(LLMs) through iterative preference learning inspired by AlphaZero, we propose to further enhance the step-wise reasoning capabilities through intrinsic self-correction to some extent. Our work leverages step-wise preference learning to enhance self-verification via reinforcement learning. We initially conduct our work through a two-stage training procedure. At the first stage, the self-correction reasoning ability of an LLM is enhanced through its own predictions, relying entirely on self-generated data within the intrinsic self-correction to some extent. At the second stage, the baseline step-wise preference learning is leveraged via the application of the enhanced self-correct policy achieved at the first stage. In the evaluation of arithmetic reasoning tasks, our approach outperforms OpenMath2-Llama3.1-8B, dart-math-mistral-7b-uniform on MATH with increases in accuracy to 71.34%(+4.18%) and 48.06%(+4.94%) and LLama-3.1-8B-Instruct, Mistral-7B-Instruct-v0.1 on GSM8K with increases in accuracy to 86.76%(+2.00%) and 38.06%(+2.28%).

推理增强自纠正大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。