arXiv:2501.10799cs.LGcs.AI2025-01被引 22

用分步二值反馈训练模型,让数学推理更可靠。

Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback

  • 结合中间步骤和最终答案的二值反馈,引导逻辑推理
  • MATH-500上准确率显著优于强基线模型
  • 适合关注可解释性与推理可信度的研究者

大型语言模型在数学推理任务中表现出色。尽管链式思维提示和自一致性采样等方法取得进展,但这些方法往往只关注最终答案正确性,而忽视推理过程的连贯性与可靠性。本文提出Step-KTO训练框架,通过中间推理步骤和最终答案的二值反馈,引导模型生成更可信的推理路径。实验表明,在复杂数学基准测试中,Step-KTO不仅提升了最终答案的准确率,也显著改善了中间步骤的质量。例如,在MATH-500数据集上,其Pass@1准确率明显优于现有强基线模型。结果表明,将分步过程反馈融入训练,有望实现更可解释、更可靠的推理能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have recently demonstrated remarkable success in mathematical reasoning. Despite progress in methods like chain-of-thought prompting and self-consistency sampling, these advances often focus on final correctness without ensuring that the underlying reasoning process is coherent and reliable. This paper introduces Step-KTO, a training framework that combines process-level and outcome-level binary feedback to guide LLMs toward more trustworthy reasoning trajectories. By providing binary evaluations for both the intermediate reasoning steps and the final answer, Step-KTO encourages the model to adhere to logical progressions rather than relying on superficial shortcuts. Our experiments on challenging mathematical benchmarks show that Step-KTO significantly improves both final answer accuracy and the quality of intermediate reasoning steps. For example, on the MATH-500 dataset, Step-KTO achieves a notable improvement in Pass@1 accuracy over strong baselines. These results highlight the promise of integrating stepwise process feedback into LLM training, paving the way toward more interpretable and dependable reasoning capabilities.

数学推理训练框架可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。