arXiv:2607.05199cs.AI2026-07

小模型物理推理出错?用分步反馈自动修正错误。

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

论文配图:Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
图 1 · 摘自论文原文
  • 通过分步奖励机制识别首错,生成针对性修正反馈。
  • 计算错误率从56.9%降至23.5%,理解错误减少至12.0%。
  • 无需人工标注,适合训练小型语言模型的自主纠错。

小语言模型在物理推理中存在结构性失败:任一步骤的错误会逐级传播,导致后续推断全盘出错。受限的领域知识、多步推导中的幻觉现象以及分布敏感性加剧了这一问题。本文提出一种分步奖励框架,可识别首个推理错误,生成针对性结构化反馈,并利用带KL正则化的策略梯度训练模型进行修正,且不依赖真实解作为生成目标。该方法无需偏好数据构建,外部验证器仅在训练阶段使用。在五个物理基准测试中,该框架相较CoT提示提升17-20%准确率,优于最强基线10-16%;计算错误率由56.9%降至23.5%,理解错误由22.3%降至12.0%;概念错误虽从89.7%降至68.7%,但仍是最难解决的失败模式。

原文摘要 · Abstract (English)

Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucination under multi-step derivation, and distributional sensitivity compound this failure. We propose a step-level reward framework that identifies the first reasoning error, generates targeted structured feedback, and trains the model to revise its solution via policy gradient with KL regularization, without exposing it to ground truth solutions as generation targets. Unlike annotation-dependent step-level methods, no preference data construction is required and the external verifier operates exclusively at training time. Across five physics benchmarks, our framework delivers accuracy gains of 17-20% over CoT prompting and 10-16% over the strongest baseline, reduces calculation errors from 56.9% to 23.5%, and reduces miscomprehension errors from 22.3% to 12.0% in the best observed cases. Conceptual errors reduce from 89.7% to 68.7%, yet persist as the hardest failure mode across all conditions.

物理推理小模型纠错强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。