系统梳理大模型数学推理研究进展与挑战
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

- 构建数学数据集统一分类体系,区分训练与评测数据
- 分析工具调用、验证器引导等策略对推理效果的影响
- 揭示当前评估指标缺陷,指出模型推理可信度不足问题
数学推理在教育、科学和工业领域至关重要,是评估人工智能系统的重要标准。随着大语言模型(LLMs)推理能力提升,理解其数学推理表现变得愈发关键。本文综述了近期基于大模型的数学推理研究,系统分析了数据集、架构、训练策略和评估协议。涵盖约120篇同行评审论文与预印本,提出数学数据集统一分类体系,区分预训练语料、监督微调资源与评估基准,按推理复杂度分级。系统分析推理架构与训练策略,包括工具集成、验证器引导推理和参数高效适配,评估其对推理鲁棒性与泛化能力的影响。比较现有评估指标,揭示最终答案准确率与过程级推理验证之间的差距。综合分析识别出重复出现的失败模式,如推理忠实度问题、基准偏差和泛化局限,并提出改进符号接地、评估可靠性及构建更稳健可信的大模型推理系统的重点方向。
原文摘要 · Abstract (English)
Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems. As Large Language Models (LLMs) improve their reasoning capabilities, understanding how well they perform mathematical reasoning has become increasingly important. This survey synthesizes recent advancements in mathematical reasoning with LLMs through a structured analysis of datasets, architectures, training strategies, and evaluation protocols. Our systematic review encompasses approximately 120 peer-reviewed studies and preprints, examining the evolution of this research area and providing a unified analytical framework to understand current progress and limitations. Our study particularly introduces a unified taxonomy of mathematical datasets, distinguishing between pretraining corpora, supervised fine-tuning resources, and evaluation benchmarks across varying levels of reasoning complexity. A systematic analysis of reasoning architectures and training strategies, including tool integration, verifier-guided reasoning, and parameter-efficient adaptation, is presented to assess their effects on reasoning robustness and generalization. Moreover, a comparative evaluation of existing metrics highlights the gap between final-answer accuracy and process-level reasoning verification. By synthesizing insights across these areas, our analysis identifies recurring failure modes, such as reasoning faithfulness issues, benchmark biases, and generalization limitations, and outlines key research directions toward improving symbolic grounding, evaluation reliability, and the development of more robust and trustworthy LLM-based reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。