arXiv:2504.06122cs.AI2025-04被引 26

用持续训练和强化学习提升形式化推理模型,实现最先进效果

Leanabell-Prover: Posttraining Scaling in Formal Reasoning

  • 通过混合数据集持续训练,融入人类推理与假设优化行为
  • 在MiniF2F上达59.8%的pass@32,超越现有模型
  • 适合形式化验证、AI数学证明方向的研究者参考

近期基于大语言模型的自动化定理证明(ATP)进展展示了使用Lean 4代码进行形式化推理的潜力。然而,与OpenAI O1/O3及Deepseek R1在自然语言推理中展现的后训练扩展效应相比,ATP尚未实现类似突破。本文系统研究了ATP的后训练过程,旨在使其与自然语言推理模型的突破同步。首先,我们利用包含大量命题-证明对的混合数据集,对现有ATP模型进行持续训练,并加入模拟人类推理与假设修正的认知行为数据。其次,采用由Lean 4编译器返回的结果奖励进行强化学习。通过设计的持续训练与强化学习流程,我们成功提升了包括DeepSeek-Prover-v1.5和Goedel-Prover在内的多个现有形式化证明器,在全证明生成任务中达到当前最优表现。例如,在MiniF2F数据集上实现59.8%的通过率(pass@32)。该项目仍在持续更新,将逐步发布数据与训练细节。

原文摘要 · Abstract (English)

Recent advances in automated theorem proving (ATP) through LLMs have highlighted the potential of formal reasoning with Lean 4 codes. However, ATP has not yet be revolutionized by the recent posttraining scaling as demonstrated by Open AI O1/O3 and Deepseek R1. In this work, we investigate the entire posttraining of ATP, aiming to align it with breakthroughs in reasoning models in natural languages. To begin, we continual train current ATP models with a hybrid dataset, which consists of numerous statement-proof pairs, and additional data aimed at incorporating cognitive behaviors that emulate human reasoning and hypothesis refinement. Next, we explore reinforcement learning with the use of outcome reward returned by Lean 4 compiler. Through our designed continual training and reinforcement learning processes, we have successfully improved existing formal provers, including both DeepSeek-Prover-v1.5 and Goedel-Prover, achieving state-of-the-art performance in the field of whole-proof generation. For example, we achieve a 59.8% pass rate (pass@32) on MiniF2F. This is an on-going project and we will progressively update our findings, release our data and training details.

形式化推理强化学习Lean 4定理证明

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。