修复数学推理中的逻辑断层,让大模型更懂一步步思考。
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
- 设计新任务自动识别推理断层并补全缺失步骤
- 在NuminaMath上提升5.87%,优于原始数据训练
- 适合改进数学推理模型,尤其对微调和强化学习有帮助
大型语言模型在数学推理任务中依赖链式思维(CoT)取得显著进展,但现有数学CoT数据集常因专家省略中间步骤导致思维断层,影响模型学习与泛化。本文提出CoT思维断层修复任务,旨在自动检测断层并生成缺失的推理步骤以恢复完整性与连贯性。为此,基于结构化的ScaleQuestMath数据集构建了专用训练数据集ScaleQM+,并训练出CoT-Bridge模型完成断层修复。在多个数学推理基准上的实验表明,使用修复后数据微调的模型性能持续优于原始数据训练的模型,最高提升达+5.87%(在NuminaMath上)。该方法还显著增强蒸馏数据效果(+3.02%),并为强化学习提供更好起点(+3.1%),可无缝集成至现有优化流程。此外,CoT-Bridge在跨领域逻辑推理任务中也展现更强泛化能力,证实提升推理完整性具有广泛收益。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from Thought Leaps due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called ScaleQM+, based on the structured ScaleQuestMath dataset, and trained CoT-Bridge to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87% on NuminaMath. Our approach effectively enhances distilled data (+3.02%) and provides better starting points for reinforcement learning (+3.1%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrate improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。