通过分步强化推理,提升大模型自我纠错能力。
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

- 分两阶段训练:先优化分步推理,再显式学习自检自纠。
- 在多个模型和数据集上,纠错频率与效果均显著提升。
- 引入教师引导的解释性理由,提供更强纠错信号。
实现有效的自我纠错(即模型自主验证并修正自身错误)仍是大语言模型的核心挑战。本文提出基于强化学习的两阶段框架SFS-DPO,实现分步级的自我验证与纠错。第一阶段通过分步偏好优化强化分步推理能力,第二阶段显式训练模型进行自我验证与修正。进一步提出教师辅助变体SFS-DPO-R,利用解释性理由增强错误验证信号,提供更优纠正指引。跨多个大模型的域内与域外评估表明,SFS-DPO与SFS-DPO-R持续优于现有分步训练基线,分析显示其在自我纠错频率与有效性方面均有提升,凸显强化分步推理对鲁棒性能的关键作用。
原文摘要 · Abstract (English)
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose Self-Fix Step-DPO (SFS-DPO), a reinforcement learning based, two-stage framework for step-level self-verification and self-correction. The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct. We further introduce a teacher-assisted variant, SFS-DPO-R, which incorporates explanatory rationales for error verification to provide stronger corrective signals. Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines. Our analysis further reveals improvements in self-correction frequency and effectiveness, highlighting the importance of strengthening step-level reasoning for robust performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。