arXiv:2506.08446cs.AIcs.CL2025-06综述被引 77

综述大模型数学推理能力的演进与挑战

A Survey on Large Language Models for Mathematical Reasoning

  • 分理解与作答两阶段,分析模型如何提升数学能力
  • 从直接预测到逐步推理,关键方法包括提示工程与强化学习
  • 适合关注大模型推理能力的研究者和应用开发者

数学推理长期是人工智能的核心挑战之一。近年来,大语言模型在该领域取得显著进展。本文从两个高层次认知阶段审视模型的数学推理能力:一是通过多样预训练策略实现数学理解的“理解”阶段;二是从直接预测演变为逐步推理(链式思维,CoT)的“答案生成”阶段。文章综述了提升数学推理的方法,涵盖无需训练的提示技巧、监督微调与强化学习等微调策略,并讨论了扩展链式思维与“测试时缩放”等新方向。尽管进展显著,模型在能力、效率和泛化方面仍面临根本性挑战。为此,本文强调了先进预训练、知识增强、形式化推理框架以及基于原则的学习范式等有前景的研究方向。本综述旨在为希望提升大模型推理能力或将其应用于其他领域的研究者提供参考。

原文摘要 · Abstract (English)

Mathematical reasoning has long represented one of the most fundamental and challenging frontiers in artificial intelligence research. In recent years, large language models (LLMs) have achieved significant advances in this area. This survey examines the development of mathematical reasoning abilities in LLMs through two high-level cognitive phases: comprehension, where models gain mathematical understanding via diverse pretraining strategies, and answer generation, which has progressed from direct prediction to step-by-step Chain-of-Thought (CoT) reasoning. We review methods for enhancing mathematical reasoning, ranging from training-free prompting to fine-tuning approaches such as supervised fine-tuning and reinforcement learning, and discuss recent work on extended CoT and "test-time scaling". Despite notable progress, fundamental challenges remain in terms of capacity, efficiency, and generalization. To address these issues, we highlight promising research directions, including advanced pretraining and knowledge augmentation techniques, formal reasoning frameworks, and meta-generalization through principled learning paradigms. This survey tries to provide some insights for researchers interested in enhancing reasoning capabilities of LLMs and for those seeking to apply these techniques to other domains.

大模型数学推理链式思维综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。