让语言和代码思维相互提升,显著增强数学推理能力
Parrot: A Training Pipeline Enhances Both Program CoT and Natural Language CoT for Reasoning
- 设计三阶段任务联合训练语言与代码思维链
- 使用辅助奖励缓解代码思维链的稀疏奖励问题
- 在MathQA上使语言思维链性能提升超21%
自然语言思维链(N-CoT)和程序思维链(P-CoT)是大模型解决数学推理问题的两大范式。现有研究多聚焦单向增强:如用P-CoT提升N-CoT,或反之。本文旨在充分挖掘两者潜力,实现双向协同增强。我们分析了两种范式的错误类型,提出Parrot训练流水线:1)三个目标导向子任务整合顺序生成语言与代码思维链;2)采用子任务混合训练策略促进自然语言语义迁移;3)设计转换后的N-CoT辅助奖励,缓解P-CoT优化中的稀疏奖励问题。大量实验表明,Parrot显著提升N-CoT与P-CoT性能,尤其在N-CoT上表现突出。使用Parrot SFT后,LLaMA2和CodeLLaMA在MathQA上的性能分别较强化学习基线提升+21.87和+21.48,且无需资源密集型训练。
原文摘要 · Abstract (English)
Natural language chain-of-thought (N-CoT) and Program chain-of-thought (P-CoT) have emerged as two primary paradigms for large language models (LLMs) to solve mathematical reasoning problems. Current research typically endeavors to achieve unidirectional enhancement: P-CoT enhanced N-CoT or N-CoT enhanced P-CoT. In this paper, we seek to fully unleash the two paradigms' strengths for mutual enhancement and ultimately achieve simultaneous improvements. We conduct a detailed analysis of the error types across two paradigms, based on which we propose Parrot, a novel training pipeline for mathematical problems: 1) Three target-designed subtasks integrate sequential P-CoT and N-CoT generation. 2) A subtask hybrid training strategy to facilitate natural language semantic transferability. 3) The converted N-CoT auxiliary reward is designed to alleviate the sparse rewards in P-CoT optimization. Extensive experiments demonstrate that Parrot significantly enhances both the performance of N-CoT and P-CoT, especially on N-CoT. Using Parrot SFT, the N-CoT performance of LLaMA2 and CodeLLaMA achieve gains of +21.87 and +21.48 on MathQA over the RL baseline, which is resource-intensive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。