解决数学推理中模型因步骤多而得分偏高的问题
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
- 用反事实分析设计去偏框架,让评分不受步骤长度影响
- 在MATH500和GSM-Plus上提升步骤选择准确率,输出更简洁
- 适合关注推理效率与公平评估的研究者与工程师
过程奖励模型(PRMs)在评估和引导大语言模型的多步推理中起关键作用,尤其在数学问题求解中。然而,我们发现现有PRMs普遍存在长度偏差:即使语义内容和逻辑有效性不变,更长的推理步骤也往往获得更高评分。这一偏差削弱了奖励预测的可靠性,并导致推理时输出过度冗长。为此,我们提出CoLD(Counterfactually-Guided Length Debiasing),一个统一框架,通过三个组件缓解长度偏差:显式的长度惩罚调整、学习到的偏差估计器以捕捉虚假的长度相关信号,以及联合训练策略,强制奖励预测具备长度不变性。该方法基于反事实推理并受因果图分析指导。在MATH500和GSM-Plus上的大量实验表明,CoLD提升了步骤选择的准确性,并促进更简洁、逻辑有效的推理。此外,它持续改善下游强化学习性能,并在不同领域间具有泛化能力,证明了其强大的通用性。
原文摘要 · Abstract (English)
Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. However, we identify a pervasive length bias in existing PRMs: they tend to assign higher scores to longer reasoning steps, even when the semantic content and logical validity are unchanged. This bias undermines the reliability of reward predictions and leads to overly verbose outputs during inference. To address this issue, we propose CoLD(Counterfactually-Guided Length Debiasing), a unified framework that mitigates length bias through three components: an explicit length-penalty adjustment, a learned bias estimator trained to capture spurious length-related signals, and a joint training strategy that enforces length-invariance in reward predictions. Our approach is grounded in counterfactual reasoning and informed by causal graph analysis. Extensive experiments on MATH500 and GSM-Plus show that CoLD improves accuracy in step selection, and encourages more concise, logically valid reasoning. Furthermore, it consistently improves downstream RL performance and generalizes across domains by mitigating length bias, demonstrating CoLD's strong generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。