用一阶逻辑证明提升大模型数学推理能力,解决多步推导易出错问题。
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving

- 引入自适应策略增强生成多样性与合理性,提升推理路径质量。
- 在447个定理数据集上,性能提升0.6%至6.4%,显著改善多步推导准确率。
- 适用于需要严谨逻辑推导的数学证明场景,尤其适合高阶推理研究者。
大语言模型(LLMs)在基础一阶逻辑(FOL)推理方面展现出潜力,但在涉及多步FOL推导的复杂数学推理中仍表现不足。例如,Deepseek-Prover-V2-7B在我们提出的定理证明数据集上的准确率仅为4.2%。这主要源于对多样化证明策略探索不足,以及早期推理错误导致整体证明失败。为此,我们提出DREAM——一种自适应方案,通过轴心驱动的策略多样化机制促进多样推理路径,并引入子命题错误反馈机制帮助模型反思和修正证明过程。本工作首次在FOL定理证明层面推动大模型数学推理进展,提出新型推理阶段解决方案,使性能提升0.6%至6.4%,并构建了包含447个定理的Lean 4格式评估数据集。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas. However, their effectiveness in complex mathematical reasoning involving multi-step FOL deductions is still under-researched. While LLMs perform competitively on established mathematical reasoning benchmarks, they struggle with multi-step FOL tasks, as demonstrated by Deepseek-Prover-V2-7B's low accuracy (4.2%) on our proposed theorem proving dataset. This issue arises from the limited exploration of diverse proof strategies and the potential for early reasoning mistakes to undermine entire proofs. To address these issues, we propose DREAM, a self-adaptive solution that enhances the Diversity and REAsonability of LLMs' generation strategies. DREAM incorporates an Axiom-Driven Strategy Diversification mechanism to promote varied strategic outcomes and a Sub-Proposition Error Feedback to help LLMs reflect on and correct their proofs. Our contributions include pioneering advancements in LLMs' mathematical reasoning through FOL theorem proving, introducing a novel inference stage solution that improves performance by 0.6% to 6.4%, and providing a curated dataset of 447 mathematical theorems in Lean 4 format for evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。