通过自适应过程优化,让小模型推理更准更快
SAPO: Self-Adaptive Process Optimization Makes Small Reasoners Stronger
- 基于错误相关负波机制,动态调整推理过程
- 在数学与代码任务上超越多数现有自进化方法
- 适合研究小模型自我改进与验证器设计的学者
现有自演化方法忽略细粒度推理步骤的影响,导致推理器与验证器之间存在差距。蒙特卡洛(MC)过程监督的计算低效进一步加剧了该问题。受错误相关负波(ERN)启发——推理器可在错误决策后定位错误并快速调整,我们提出一种自适应过程优化(SAPO)方法,用于小型语言模型(SLMs)的自我提升。SAPO通过主动最小化推理器-验证器差距,自适应且高效地引入过程监督信号,而非依赖低效的MC估计。大量实验表明,该方法在数学和代码两类挑战性任务上优于多数现有自演化方法。此外,为深入探究SAPO对验证器性能的影响,本文还构建了两个新基准,分别用于数学与编码任务中的过程奖励模型评估。
原文摘要 · Abstract (English)
Existing self-evolution methods overlook the influence of fine-grained reasoning steps, which leads to the reasoner-verifier gap. The computational inefficiency of Monte Carlo (MC) process supervision further exacerbates the difficulty in mitigating the gap. Motivated by the Error-Related Negativity (ERN), which the reasoner can localize error following incorrect decisions, guiding rapid adjustments, we propose a Self-Adaptive Process Optimization (SAPO) method for self-improvement in Small Language Models (SLMs). SAPO adaptively and efficiently introduces process supervision signals by actively minimizing the reasoner-verifier gap rather than relying on inefficient MC estimations. Extensive experiments demonstrate that the proposed method outperforms most existing self-evolution methods on two challenging task types: mathematics and code. Additionally, to further investigate SAPO's impact on verifier performance, this work introduces two new benchmarks for process reward models in both mathematical and coding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。