arXiv:2507.20199cs.AI2025-07被引 11

让AI一步步思考并验证,用工具交互提升数学定理证明能力

StepFun-Prover Preview: Let's Think and Verify Step by Step

  • 通过强化学习融合工具交互,实现逐步推理与反馈修正
  • 在miniF2F-test上达70.0%的pass@1成功率,采样极少
  • 适合自动化定理证明与数学AI助手研发者参考

我们提出StepFun-Prover Preview,一种通过工具集成推理实现形式化定理证明的大语言模型。该模型采用结合工具交互的强化学习流程,在生成Lean 4证明时表现出色,仅需少量采样即可达到高精度。其方法模拟人类解题策略,通过实时环境反馈迭代优化证明过程。在miniF2F-test基准测试中,StepFun-Prover达到70.0%的pass@1成功率。除了提升基准表现,我们还构建了端到端训练框架,为开发工具集成推理模型提供新路径,推动自动化定理证明与数学AI助手的发展。

原文摘要 · Abstract (English)

We present StepFun-Prover Preview, a large language model designed for formal theorem proving through tool-integrated reasoning. Using a reinforcement learning pipeline that incorporates tool-based interactions, StepFun-Prover can achieve strong performance in generating Lean 4 proofs with minimal sampling. Our approach enables the model to emulate human-like problem-solving strategies by iteratively refining proofs based on real-time environment feedback. On the miniF2F-test benchmark, StepFun-Prover achieves a pass@1 success rate of $70.0\%$. Beyond advancing benchmark performance, we introduce an end-to-end training framework for developing tool-integrated reasoning models, offering a promising direction for automated theorem proving and Math AI assistant.

定理证明强化学习数学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。