让大模型自动发现并修正推理错误,提升数学解题能力。
S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners

- 在每一步推理中自动检测错误并修正,无需外部干预。
- 在GSM8K和MATH数据集上显著提升准确率,最高达+12.3%。
- 适用于各类基础大模型,适合需要高可靠推理的场景。
自校正是一种能激发大语言模型(LLMs)推理潜力的新方法,通过在推理过程中检测并修正错误来实现。然而,现有工作并未将自校正视为大模型的自发内在能力,而是依赖事后生成、引入外部知识或多模型协作等手段。本文提出一系列名为S^3cMath的数学大模型,具备自发的逐步自校正能力,可识别推理过程中的潜在错误并即时修正,从而生成更可靠的解答。我们设计了一种基于步骤采样的方法,构建了用于训练的逐步自校正数据;并采用相应的训练策略,使大模型获得这种自发纠错能力。实验表明,该方法在多种基础大模型上均有效,在GSM8K和MATH等数学基准测试中持续取得显著提升。据我们所知,这是首次在数学推理中引入大模型自发逐步自校正能力。
原文摘要 · Abstract (English)
Self-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs). It involves detecting and correcting errors during the inference process when LLMs solve reasoning problems. However, recent works do not regard self-correction as a spontaneous and intrinsic capability of LLMs. Instead, such correction is achieved through post-hoc generation, external knowledge introduction, multi-model collaboration, and similar techniques. In this paper, we propose a series of mathematical LLMs called S$^3$c-Math, which are able to perform Spontaneous Step-level Self-correction for Mathematical reasoning. This capability helps LLMs to recognize whether their ongoing inference tends to contain errors and simultaneously correct these errors to produce a more reliable response. We proposed a method, which employs a step-level sampling approach to construct step-wise self-correction data for achieving such ability. Additionally, we implement a training strategy that uses above constructed data to equip LLMs with spontaneous step-level self-correction capacities. Our data and methods have been demonstrated to be effective across various foundation LLMs, consistently showing significant progress in evaluations on GSM8K, MATH, and other mathematical benchmarks. To the best of our knowledge, we are the first to introduce the spontaneous step-level self-correction ability of LLMs in mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。