让物理专家低成本生成可验证的正式证明,突破了科学领域自动形式化的瓶颈。
FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in Lean

- 构建人机协作的智能体流程,专家仅需简单交互即可生成正确形式化代码。
- 创建200个大学级物理问题的正式表示数据集,证明完全正确且复杂度更高。
- 首次揭示符号坍缩、抽象提升等语义漂移现象,指导未来模型改进。
将非形式化数学推理转化为可验证代码是大语言模型面临的重大挑战。在物理等科学领域,特定符号体系(如狄拉克符号、向量微积分)进一步加剧了形式化难度,现有大模型与智能体方法尚未有效应对。为此,我们提出 FormalScience:一种无需领域专家精通形式语言的通用型人机协同智能体流程,使单名领域专家以低经济成本生成语法正确、语义对齐的形式化证明。应用于物理领域,我们构建了 FormalPhysics——一个包含200个大学水平(LaTeX)物理问题及解答(主要为量子力学与电磁学)及其 Lean4 形式化表示的数据集。相较于现有形式数学基准,FormalPhysics 实现了100%的形式有效性,并表现出更高的命题复杂度。我们通过零样本提示、自精炼结合错误反馈以及新型多阶段智能体方法,在该数据集上评估开源与专有模型在命题自动形式化任务中的表现,并系统分析了现代基于LLM方法在自动形式化中的局限性。首次从符号坍缩、抽象提升等概念出发,系统刻画了物理自动形式化中的语义漂移现象,揭示当完全语义保真不可达时,形式语言实际验证的内容。我们开源代码库及基于交互界面的 FormalScience 系统,支持科学领域(包括但不限于物理)的自动形式化与定理证明。
原文摘要 · Abstract (English)
Formalising informal mathematical reasoning into formally verifiable code is a significant challenge for large language models. In scientific fields such as physics, domain-specific machinery (\textit{e.g.} Dirac notation, vector calculus) imposes additional formalisation challenges that modern LLMs and agentic approaches have yet to tackle. To aid autoformalisation in scientific domains, we present FormalScience; a domain-agnostic human-in-the-loop agentic pipeline that enables a single domain expert (without deep formal language experience) to produce \textit{syntactically correct} and \textit{semantically aligned} formal proofs of informal reasoning for low economic cost. Applying FormalScience to physics, we construct FormalPhysics, a dataset of 200 university-level (LaTeX) physics problems and solutions (primarily quantum mechanics and electromagnetism), along with their Lean4 formal representations. Compared to existing formal math benchmarks, FormalPhysics achieves perfect formal validity and exhibits greater statement complexity. We evaluate open-source models and proprietary systems on a statement autoformalisation task on our dataset via zero-shot prompting, self-refinement with error feedback, and a novel multi-stage agentic approach, and explore autoformalisation limitations in modern LLM-based approaches. We provide the first systematic characterisation of semantic drift in physics autoformalisation in terms of concepts such as notational collapse and abstraction elevation which reveals what formal language verifies when full semantic preservation is unattainable. We release the codebase together with an interactive UI-based FormalScience system which facilitates autoformalisation and theorem proving in scientific domains beyond physics.https://github.com/jmeadows17/formal-science
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。