用检索增强的LLM提升Lean 4在大学数学定理证明中的表现
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
- 基于微调LLM与检索系统,构建分步推理的Lean 4定理证明器
- 在ProofNet上达23.7%成功率,在新基准FATE-M上达56.7%领先水平
- 开源数据管道与交互环境助力高质量训练数据生成
当前形式化定理证明器在中学和竞赛级数学上取得显著进展,但难以推广到更高级数学。本文提出REAL-Prover,一个面向Lean 4的开源分步定理证明器,旨在突破这一界限。该证明器基于微调的大语言模型REAL-Prover-v1,并集成检索系统Leansearch-PS,显著提升解决大学级数学问题的能力。为训练REAL-Prover-v1,我们开发了HERALD-AF数据提取流水线,将自然语言数学题转为形式化陈述,并构建了新的开源Lean 4交互环境Jixia-interactive以促进合成数据收集。实验显示,仅通过监督微调,其在ProofNet数据集上达到23.7%的通过率(Pass@64),与现有最先进模型相当。为进一步评估,我们引入聚焦代数问题的新基准FATE-M,REAL-Prover在此达到56.7%的通过率(Pass@64),刷新SOTA记录。
原文摘要 · Abstract (English)
Nowadays, formal theorem provers have made monumental progress on high-school and competition-level mathematics, but few of them generalize to more advanced mathematics. In this paper, we present REAL-Prover, a new open-source stepwise theorem prover for Lean 4 to push this boundary. This prover, based on our fine-tuned large language model (REAL-Prover-v1) and integrated with a retrieval system (Leansearch-PS), notably boosts performance on solving college-level mathematics problems. To train REAL-Prover-v1, we developed HERALD-AF, a data extraction pipeline that converts natural language math problems into formal statements, and a new open-source Lean 4 interactive environment (Jixia-interactive) to facilitate synthesis data collection. In our experiments, our prover using only supervised fine-tune achieves competitive results with a 23.7% success rate (Pass@64) on the ProofNet dataset-comparable to state-of-the-art (SOTA) models. To further evaluate our approach, we introduce FATE-M, a new benchmark focused on algebraic problems, where our prover achieves a SOTA success rate of 56.7% (Pass@64).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。