首个面向物理领域的形式化定理证明系统,用少量数据显著提升推理能力。
PhysProver: Advancing Automatic Theorem Proving for Physics
- 构建物理专用数据集PhysLeanData,结合猜想生成与真实定理
- 仅用约5000样本训练,多领域平均提升2.4%,跨域测试增1.3%
- 开源模型与数据,推动形式化推理向物理等新领域扩展
形式化语言与大模型的结合已深刻影响数学与计算机科学,为定理证明提供严谨基础。尽管近期在数学推理领域取得进展,但物理领域的形式化推理仍被忽视。本文提出迄今首个面向物理领域的形式化定理证明方法。我们构建了专用数据集PhysLeanData,包含从PhysLean中采样的定理及基于猜想的形式化生成数据。训练中采用DeepSeek-Prover-V2-7B作为基座模型,并应用可验证奖励的强化学习(RLVR)进行优化。实验表明,仅使用约5,000个训练样本,PhysProver在多个子领域实现2.4%的整体性能提升。此外,在MiniF2F-Test基准上观察到1.3%的增益,表明其具备跨域泛化能力,同时增强了形式化数学推理能力。结果验证了方法的有效性与高效性,为将形式化证明器拓展至数学以外领域提供了范式。为促进研究,我们将公开数据集与模型。
原文摘要 · Abstract (English)
The combination of verifiable languages and LLMs has significantly influenced both the mathematical and computer science communities because it provides a rigorous foundation for theorem proving. Recent advancements in the field provide foundation models and sophisticated agentic systems pushing the boundaries of formal mathematical reasoning to approach the natural language capability of LLMs. However, little attention has been given to the formal physics reasoning, which also heavily relies on similar problem-solving and theorem-proving frameworks. To solve this problem, this paper presents, to the best of our knowledge, the first approach to enhance formal theorem proving in the physics domain. We compose a dedicated dataset PhysLeanData for the task. It is composed of theorems sampled from PhysLean and data generated by a conjecture-based formal data generation pipeline. In the training pipeline, we leverage DeepSeek-Prover-V2-7B, a strong open-source mathematical theorem prover, and apply Reinforcement Learning with Verifiable Rewards (RLVR) to train our model PhysProver. Comprehensive experiments demonstrate that, using only $\sim$5K training samples, PhysProver achieves an overall 2.4\% improvement in multiple sub-domains. Furthermore, after formal physics training, we observe 1.3\% gains on the MiniF2F-Test benchmark, which indicates non-trivial generalization beyond physics domains and enhancement for formal math capability as well. The results highlight the effectiveness and efficiency of our approach, which provides a paradigm for extending formal provers outside mathematical domains. To foster further research, we will release both our dataset and model to the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。