长思维链训练能提升模型跨领域推理能力,短链训练效果有限。
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
- 用长思维链和规则强化学习训练数学求解
- 持续预训练可泛化到8个通用推理任务
- 适合想提升模型泛化推理能力的研究者
大语言模型的数学问题求解能力日益受到关注。现有研究多聚焦于构建专用模型,但数学求解能力能否泛化至其他推理任务仍不明确。本文通过实证研究评估了连续预训练、指令微调及基于规则的强化学习在多种数据源(含短链与长链思维链样本)上的表现。在5个数学和8个通用推理基准上的评估显示,对数学文本进行持续预训练可在一定程度上泛化至通用推理任务;而仅在常规短链样本上进行指令微调则效果有限,甚至损害泛化性能。值得注意的是,使用长思维链样本训练并结合规则强化学习显著提升了泛化能力,使模型推理过程延伸至其他领域。结果表明,传统短链训练难以实现稳健泛化,而长思维链结合自我反思的新范式,为通过专精领域学习提升通用推理能力提供了前景。
原文摘要 · Abstract (English)
There has been a growing interest in enhancing the mathematical problem-solving (MPS) capabilities of large language models. While the majority of research efforts concentrate on creating specialized models to solve mathematical problems, it remains unknown how learning mathematical problem-solving generalizes to help develop other reasoning abilities. In this paper, we present an empirical investigation into the generalization potential of various MPS training approaches, such as continual pretraining, instruction tuning, and rule-based reinforcement learning across various data sources, including both short and long chain-of-thought (CoT) samples. Evaluation on 5 mathematical and 8 general reasoning benchmarks show that continual pretraining on math text is able to generalize to general reasoning tasks to some extent. In constrast, instruction tuning on conventional, short MPS samples provides limited benefits and, in many cases, even impairs generalization performance. Notably, training with long CoT responses for MPS samples and incorporating rule-based reinforcement learning on MPS queries exhibit distinct behavior, significantly enhancing generalization by extending the model's reasoning processes into other domains. These results suggest that traditional approaches to learning MPS with short reasoning chains largely fail to achieve robust generalization. However, the emerging paradigm of longer reasoning chains, coupled with self-reflection, offers a promising direction for improving generalized reasoning abilities through learning from specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。