arXiv:2510.25427cs.CLcs.AI2025-10EMNLP被引 8

构建真实数学定理证明评估集,揭示大模型在实际场景中表现仍差。

RLMEval: Evaluating Research-Level Neural Theorem Proving

  • 基于真实Lean项目构建研究级定理证明评估集
  • 顶尖模型在613个定理上仅达10.3%通过率
  • 适合关注形式化数学自动推理的研究者

尽管在精心设计的基准测试中表现优异,大型语言模型(LLMs)在研究级神经定理证明和证明自动化形式化方面的实际影响仍然有限。我们提出RLMEval,一个针对这些任务的评估套件,聚焦于来自真实世界Lean形式化项目的高水平数学问题。RLMEval通过利用真实的Lean Blueprint形式化项目,对神经定理证明和证明自动化形式化在具有挑战性的研究级定理上的表现进行评估。我们在包含613个定理的6个Lean项目上的评估表明,现有基准上的进展无法直接迁移到更现实的场景中:最先进模型仅达到10.3%的通过率。RLMEval提供了一个新的、具有挑战性的基准,旨在引导并加速形式化数学自动化推理的发展。

原文摘要 · Abstract (English)

Despite impressive results on curated benchmarks, the practical impact of large language models (LLMs) on research-level neural theorem proving and proof autoformalization is still limited. We introduce RLMEval, an evaluation suite for these tasks, focusing on research-level mathematics from real-world Lean formalization projects. RLMEval targets the evaluation of neural theorem proving and proof autoformalization on challenging research-level theorems by leveraging real Lean Blueprint formalization projects. Our evaluation of state-of-the-art models on RLMEval, comprising 613 theorems from 6 Lean projects, reveals a significant gap: progress on existing benchmarks does not readily translate to these more realistic settings, with the best model achieving only a 10.3 % pass rate. RLMEval provides a new, challenging benchmark designed to guide and accelerate progress in automated reasoning for formal mathematics.

定理证明形式化数学评估基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。