7B模型通过三阶段训练,显著提升数学定理证明能力。
Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training
- 分三阶段训练:预训练+专家微调+强化优化,逐步增强推理能力。
- 在普特南竞赛题上解决27题(pass@32),组合数学题达24.0%成功率。
- 适合对轻量级形式化推理感兴趣的开发者和研究者。
大型语言模型在自动定理证明中展现出巨大潜力,但进展常受限于高质量、多样化的形式化语言数据稀缺。为此,我们提出Spark-Prover-X1,一个70亿参数的模型,采用三阶段训练框架,旨在激发更易获取的中小型语言模型的推理潜能。第一阶段通过大规模数学语料连续预训练,并引入新型数据任务,关键创新为“CoT增强的状态预测”任务,实现细粒度推理。第二阶段在专家迭代循环中进行监督微调,专门优化Spark-Prover-X1-7B与Spark-Formalizer-X1-7B模型。最后,针对最难题目执行定向的组相对策略优化(GRPO)以进一步提升证明能力。为支持稳健评估,尤其在真实考试题目上的表现,我们还引入ExamFormal-Bench,一个包含402个形式化问题的新基准数据集。实验表明,Spark-Prover在同类开源模型中达到当前最佳性能,在‘完整证明生成’范式下表现优异。其在高难度竞赛基准上尤为突出,于PutnamBench上成功解决27题(pass@32),在CombiBench上达成24.0%(pass@32)的成绩。本工作验证了多样化数据与渐进式训练流程对提升轻量级模型形式化推理能力的有效性。我们即将发布Spark-Prover-X1-7B、Spark-Formalizer-X1-7B及ExamFormal-Bench数据集。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown significant promise in automated theorem proving, yet progress is often constrained by the scarcity of diverse and high-quality formal language data. To address this issue, we introduce Spark-Prover-X1, a 7B parameter model trained via an three-stage framework designed to unlock the reasoning potential of more accessible and moderately-sized LLMs. The first stage infuses deep knowledge through continuous pre-training on a broad mathematical corpus, enhanced by a suite of novel data tasks. Key innovation is a "CoT-augmented state prediction" task to achieve fine-grained reasoning. The second stage employs Supervised Fine-tuning (SFT) within an expert iteration loop to specialize both the Spark-Prover-X1-7B and Spark-Formalizer-X1-7B models. Finally, a targeted round of Group Relative Policy Optimization (GRPO) is applied to sharpen the prover's capabilities on the most challenging problems. To facilitate robust evaluation, particularly on problems from real-world examinations, we also introduce ExamFormal-Bench, a new benchmark dataset of 402 formal problems. Experimental results demonstrate that Spark-Prover achieves state-of-the-art performance among similarly-sized open-source models within the "Whole-Proof Generation" paradigm. It shows exceptional performance on difficult competition benchmarks, notably solving 27 problems on PutnamBench (pass@32) and achieving 24.0\% on CombiBench (pass@32). Our work validates that this diverse training data and progressively refined training pipeline provides an effective path for enhancing the formal reasoning capabilities of lightweight LLMs. We will release both Spark-Prover-X1-7B and Spark-Formalizer-X1-7B, along with the ExamFormal-Bench dataset, in the near future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。