arXiv:2603.12744cs.LGcs.AI2026-03被引 3

测试大模型在非标准数学定义下的推理能力,发现性能下降26%。

TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?

  • 构建新基准TaoBench,从头构造分析学概念,脱离Mathlib框架
  • 相同题目在TaoBench上平均表现比Mathlib低26%,反映泛化不足
  • 适合关注数学证明模型真实适用性的研究者

自动化定理证明(ATP)基准多基于MathLib形式化问题,导致当前训练与评估严重依赖其定义框架。然而前沿数学常具探索性,依赖定制化构造,偏离标准库。本文引入TaoBench,一个源自陶哲轩《Analysis I》的本科级基准,通过从头构建核心数学概念,不依赖Mathlib定义,并混合使用从头与Mathlib构造。为公平评估,构建智能流水线自动提取可编译的独立环境;并为每个问题生成数学等价的Mathlib版本,形成配对的TaoBench-Mathlib命题。尽管顶尖ATP模型在MathLib框架下表现良好,但在等价的Tao形式下平均性能下降约26%。这表明主要瓶颈在于跨定义框架的泛化能力,而非任务难度本身。TaoBench揭示了基准性能与实际应用间的差距,为开发更契合研究数学的证明器提供了实证基础。

原文摘要 · Abstract (English)

Automated theorem proving (ATP) benchmarks largely consist of problems formalized in MathLib, so current ATP training and evaluation are heavily biased toward MathLib's definitional framework. However, frontier mathematics is often exploratory and prototype-heavy, relying on bespoke constructions that deviate from standard libraries. In this work, we evaluate the robustness of current ATP systems when applied to a novel definitional framework, specifically examining the performance gap between standard library problems and bespoke mathematical constructions. We introduce TaoBench, an undergraduate-level benchmark derived from Terence Tao's Analysis I, which formalizes analysis by constructing core mathematical concepts from scratch, without relying on standard Mathlib definitions, as well as by mixing from-scratch and MathLib constructions. For fair evaluation, we build an agentic pipeline that automatically extracts a compilable, self-contained local environment for each problem. To isolate the effect of definitional frameworks, we additionally translate every problem into a mathematically equivalent Mathlib formulation, yielding paired TaoBench-Mathlib statements for direct comparison. While state-of-the-art ATP models perform capably within the MathLib framework, performance drops by an average of roughly 26% on the definitionally equivalent Tao formulation. This indicates that the main bottleneck is limited generalization across definitional frameworks rather than task difficulty. TaoBench thus highlights a gap between benchmark performance and applicability, and provides a concrete foundation for developing and testing provers better aligned with research mathematics.

定理证明大模型数学推理泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。