让形式化验证更智能,自动推荐证明所需前提。
Premise Selection for a Lean Hammer

- 专为依赖类型理论设计的神经前提选择模型。
- 比现有方法多解决21%的证明目标,跨领域泛化能力强。
- 适合研究者和工程师快速上手形式化验证。
神经方法正在改变证明助手中的自动化推理,但将其融入实际验证工作流仍具挑战。锤子(hammer)是一种整合前提选择、翻译至外部自动定理证明器及证明重构的综合工具,用以自动化繁琐的推理步骤。本文提出LeanPremise——一种新型神经前提选择系统,并与现有的翻译和证明重构组件结合,构建出首个面向Lean证明助手的端到端领域泛化的锤子工具:LeanHammer。与现有Lean前提选择器不同,LeanPremise专为在依赖类型理论中使用锤子而训练,能够动态适应用户特定上下文,有效推荐训练数据外的库命题及用户本地定义的引理。通过全面评估,我们发现LeanPremise使LeanHammer相比现有前提选择器多解决21%的目标,并在多样领域中表现良好。本工作有助于弥合神经检索与符号推理之间的差距,使形式化验证对研究人员和实践者更易获取。
原文摘要 · Abstract (English)
Neural methods are transforming automated reasoning for proof assistants, yet integrating these advances into practical verification workflows remains challenging. A hammer is a tool that integrates premise selection, translation to external automatic theorem provers, and proof reconstruction into one overarching tool to automate tedious reasoning steps. We present LeanPremise, a novel neural premise selection system, and we combine it with existing translation and proof reconstruction components to create LeanHammer, the first end-to-end domain general hammer for the Lean proof assistant. Unlike existing Lean premise selectors, LeanPremise is specifically trained for use with a hammer in dependent type theory. It also dynamically adapts to user-specific contexts, enabling it to effectively recommend premises from libraries outside LeanPremise's training data as well as lemmas defined by the user locally. With comprehensive evaluations, we show that LeanPremise enables LeanHammer to solve 21% more goals than existing premise selectors and generalizes well to diverse domains. Our work helps bridge the gap between neural retrieval and symbolic reasoning, making formal verification more accessible to researchers and practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。