构建首个理论级逻辑形式化基准,推动计算机科学形式验证自动化。
Theory-Scale Auto-Formalization of Logics for Computer Science
- 设计半自动代理流程,结合概念图与签名规划实现理论级翻译。
- 覆盖327个教材条目、超4076条声明,最先进模型仅达20.1%准确率。
- 提供细粒度评估协议,适合研究形式化与自动化推理的学者。
自动形式化对可扩展的形式化验证至关重要,但现有工作多集中于孤立命题,而理论级自动形式化——即一致地转换数百个相互依赖的定义、引理和定理——仍面临一致性、忠实性、可扩展性和正确性的挑战。本文提出LCS-Bench,一个基于《计算机科学逻辑》的独立理论级基准。该基准通过新颖的半自动代理流水线构建,融合概念图、形式签名规划、问题追踪及反例搜索补全(sorry-filling),并经由人类专家进行忠实性审查。最终成果涵盖327个教材条目、超过4,076条Lean声明和85,000余行Lean代码。数据引擎可自动生成五个评估赛道,衡量不同维度的自动形式化与定理证明能力。我们引入新的定义等价性检查器评估协议,实现更精细、更忠实的评估。在14个模型上的广泛实验表明:(1) LCS-Bench质量高、一致且忠实;(2) 该基准极具挑战性,最先进模型在自动形式化任务上仅达20.1%准确率;(3) 分析揭示了理论级形式化的关键洞见,并指明未来研究方向。
原文摘要 · Abstract (English)
Auto-formalization is critical for scalable formal verification, but existing progress largely focuses on isolated statements, while theory-scale auto-formalization, which coherently translates hundreds of interdependent definitions, lemmas, and theorems, remains open due to challenges in consistency, faithfulness, scalability, and correctness. In this paper, we introduce LCS-Bench, a stand-alone, theory-scale benchmark based on Logics for Computer Science. LCS-Bench is built through a novel semi-automated agentic pipeline that leverages concept graphs, formal signature planning, issue tracking, sorry-filling with counter-example search, complemented by faithfulness review from human experts. The resulting artifact covers 327 textbook items, over 4,076 Lean declarations, and more than 85K lines of Lean code. The dataset supports broad evaluation through a data engine that automatically derives five tracks of evaluation benchmarks, measuring different aspects of auto-formalization and theorem-proving capabilities. We also introduce a novel evaluation protocol featuring definitional equivalence checkers, enabling more fine-grained and faithful assessment. Through extensive evaluation on 14 models, we demonstrate that (1) LCS-Bench is of high quality, consistent, and faithful; (2) the benchmark is challenging, with state-of-the-art models achieving only 20.1% on auto-formalization tasks; and (3) our analysis reveals key findings regarding theory-scale auto-formalization and suggests promising directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。