探讨大模型在数学推理中的表现与局限,揭示其与编程能力的差异。
Thinking Machines: Mathematical Reasoning in the Age of LLMs
- 对比自然语言数学与形式化数学的训练效果差异
- 发现证明生成比代码生成更易出错且不稳定
- 质疑大模型是否真具备逻辑演进能力
大型语言模型(LLMs)在结构化推理和符号任务中表现出色,尤其在编程领域成果显著。这一进展自然推动了将模型扩展至数学领域,涵盖自然语言表述的传统数学和适合自动验证的形式化数学。然而,尽管编程与证明构造存在表面相似性,形式化数学的进展却远为困难。这一差距引发对当前大模型推理本质的深层思考:包括监督反馈的作用、模型是否保持计算或演绎状态等。本文综述了当前大模型在数学推理方面的最新进展,聚焦于近期模型与评测基准,探讨三个核心问题:(i) 传统数学与形式化数学作为训练与评估领域的权衡;(ii) 为何证明生成仍比代码生成更脆弱;(iii) 大模型是否真正表征而非仅模拟逻辑状态的演化。目标并非划定绝对界限,而是厘清现有系统的能力边界,并指明未来拓展方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in structured reasoning and symbolic tasks, with coding emerging as a particularly successful application. This progress has naturally motivated efforts to extend these models to mathematics, both in its traditional form, expressed through natural-style mathematical language, and in its formalized counterpart, expressed in a symbolic syntax suitable for automatic verification. Yet, despite apparent parallels between programming and proof construction, advances in formalized mathematics have proven significantly more challenging. This gap raises fundamental questions about the nature of reasoning in current LLM architectures, the role of supervision and feedback, and the extent to which such models maintain an internal notion of computational or deductive state. In this article, we review the current state-of-the-art in mathematical reasoning with LLMs, focusing on recent models and benchmarks. We explore three central issues at the intersection of machine learning and mathematical cognition: (i) the trade-offs between traditional and formalized mathematics as training and evaluation domains; (ii) the structural and methodological reasons why proof synthesis remains more brittle than code generation; and (iii) whether LLMs genuinely represent or merely emulate a notion of evolving logical state. Our goal is not to draw rigid distinctions but to clarify the present boundaries of these systems and outline promising directions for their extension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。