融合多种推理方式,让大模型更懂数学题
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective

- 用自然语言、算法和符号三种推理协同解题
- 在定理证明上比GPT-4o高41%,算术任务超强化学习方法15%
- 适合需要跨领域数学推理的科研与教育场景
大型语言模型在数学推理方面取得显著进展,但通常依赖单一推理范式,限制了其在多样化任务中的表现。本文提出链式推理(CoR)框架,融合自然语言推理(NLR)、算法推理(AR)和符号推理(SR)三种范式,实现协同工作。CoR通过不同范式生成多个候选答案,并整合为一致的最终解答。我们设计渐进式范式训练(PPT)策略,使模型逐步掌握各推理范式,构建出CoR-Math-7B模型。实验表明,CoR-Math-7B显著优于现有最先进模型:在定理证明任务中较GPT-4o绝对提升41.0%,在MATH基准算术任务上较基于强化学习的方法提升15.0%。结果验证了该模型在数学理解上的增强能力,支持零样本跨任务泛化。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have made notable progress in mathematical reasoning, yet often rely on single-paradigm reasoning, limiting their effectiveness across diverse tasks. We introduce Chain-of-Reasoning (CoR), a novel unified framework integrating multiple reasoning paradigms--Natural Language Reasoning (NLR), Algorithmic Reasoning (AR), and Symbolic Reasoning (SR)--to enable synergistic collaboration. CoR generates multiple potential answers via different reasoning paradigms and synthesizes them into a coherent final solution. We propose a Progressive Paradigm Training (PPT) strategy for models to progressively master these paradigms, leading to CoR-Math-7B. Experimental results demonstrate that CoR-Math-7B significantly outperforms current SOTA models, achieving up to a 41.0% absolute improvement over GPT-4o in theorem proving and a 15.0% improvement over RL-based methods on the MATH benchmark in arithmetic tasks. These results show the enhanced mathematical comprehension ability of our model, enabling zero-shot generalization across tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。