构建首个大学物理形式化推理框架,推动数学证明与物理建模结合。
Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4
- 基于Lean4构建物理定理库和基准测试集,支持形式化推理。
- 顶尖模型在该任务上最高仅达35%准确率,凸显挑战性。
- 社区共建的物理基础库可提升模型性能11.75%,适合研究形式化数学的学者。
我们提出Lean4PHYS,一个面向大学物理问题的形式化推理框架。该框架包含*LeanPhysBench*——一个由大学教材与物理竞赛题手工构建并经同行评审的200条陈述组成的大学级物理推理基准。为建立形式化物理推理基础,我们还引入*PhysLib*,一个社区驱动的物理基础库,涵盖基本单位系统与关键定理。基于此框架,我们评估了主流数学类Lean4证明器及最先进的闭源大模型,结果显示最佳模型DeepSeek-Prover-V2-7B仅达到16%准确率,Claude-Sonnet-4为35%。详细分析表明,使用*PhysLib*可使模型平均性能提升11.75%。这体现了*LeanPhysBench*的高难度与*PhysLib*的有效性。据我们所知,这是首个在Lean4中构建的物理推理基准。
原文摘要 · Abstract (English)
We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. **Lean4PHYS** includes *LeanPhysBench*, a college-level benchmark for formal physics reasoning in Lean4, which contains 200 hand-crafted and peer-reviewed statements derived from university textbooks and physics competition problems. To establish a solid foundation for formal reasoning in physics, we also introduce *PhysLib*, a community-driven repository containing fundamental unit systems and theorems essential for formal physics reasoning. Based on the benchmark and Lean4 repository we composed in **Lean4PHYS**, we report baseline results using major expert Math Lean4 provers and state-of-the-art closed-source models, with the best performance of DeepSeek-Prover-V2-7B achieving only 16% and Claude-Sonnet-4 achieving 35%. We also conduct a detailed analysis showing that our *PhysLib* can achieve an average improvement of 11.75% in model performance. This demonstrates the challenging nature of our *LeanPhysBench* and the effectiveness of *PhysLib*. To the best of our knowledge, this is the first study to provide a physics benchmark in Lean4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。