研究强化学习训练如何影响数学推理的中间步骤质量
Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- 用一阶逻辑衡量推理步骤的一致性,区分有效性与局部连贯性
- 强化学习提升推理步骤连贯性,但不保证最终答案正确
- 发现局部连贯性提升未必带来有效数学推导,需谨慎解读改进声明
基于可验证奖励的强化学习(RLVR)后训练大语言模型在推理任务中已被证明能提高准确性,持续受到关注。然而,现有方法通常对所有标记一视同仁,未考虑标记级别的优势。这些方法主要依据最终答案正确性或Pass@K指标评估性能,却声称强化学习提升了推理过程。这促使我们探究强化学习对非直接激励的中间标记的影响。为此,我们在GSM8K数据集上使用GRPO算法和Qwen-2.5-0.5B模型设计实验,引入基于一阶逻辑(FOL)的推理轨迹连贯性度量,以识别轨迹中的错误。我们区分了轨迹有效性(逻辑严谨)与轨迹连贯性(无错误)。结果显示,强化学习整体提升了轨迹连贯性,尤其在基础模型失败而强化模型成功的问题上增益显著。令人惊讶的是,强化学习提升了局部连贯性,但并未必然产生有效或正确的解。这揭示了关键区别:推理步骤的连贯性提升并不保证最终答案正确。我们主张,关于强化学习改善推理的说法需谨慎审视,因为其可能仅源于轨迹连贯性提升,而无法转化为完整有效的数学证明。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR)-based post-training of Large Language Models (LLMs) has been shown to improve accuracy on reasoning tasks and continues to attract significant attention. Existing RLVR methods, however, typically treat all tokens uniformly without accounting for token-level advantages. These methods primarily evaluate performance based on final answer correctness or Pass@K accuracy, and yet make claims about RL post-training leading to improved reasoning traces. This motivates our investigation into the effect of RL post-training on intermediate tokens which are not directly incentivized. To study this, we design an experimental setup using the GRPO algorithm with Qwen-2.5-0.5B model on the GSM8K dataset. We introduce trace coherence, a First-Order Logic (FOL)-based measure to capture the consistency of reasoning steps by identifying errors in the traces. We distinguish between trace validity and trace coherence, noting that the former implies logical soundness while the latter measures local coherence via lack of errors. Our results show that RL post-training overall improves trace coherence with the most significant gains on problems where the base model fails but the RL model succeeds. Surprisingly, RL enhances local coherence without necessarily producing valid or correct solutions. This highlights a crucial distinction: improved local coherence in reasoning steps does not guarantee final answer correctness. We argue that claims of improved reasoning via RL must be examined with care, as these may be based on improved trace coherence, which may not translate into fully valid mathematical proofs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。