arXiv:2605.14061cs.AIcs.LG2026-05被引 1

首个面向研究生数学的自动形式化基准,涵盖5万+定理与依赖关系。

MathAtlas: A Benchmark for Autoformalization in the Wild

论文配图:MathAtlas: A Benchmark for Autoformalization in the Wild
图 1 · 摘自论文原文
  • 从103本研究生教材中提取52,000条数学内容,构建带依赖图谱的基准集。
  • 现有模型在定理形式化上最高仅9.8%正确率,深度依赖项下降至2.6%。
  • 适合研究形式化推理、数学AI和依赖建模的学者使用。

当前自动形式化基准主要聚焦奥数或本科数学,而研究生及研究级数学仍缺乏探索。本文提出MathAtlas,首个大规模、真实场景下的研究生级别数学自动形式化基准,包含约52,000条定理、定义、习题、例题与证明,数据源自103本研究生数学教材。该基准还构建了约178,000条数学依赖关系的依赖图谱,是首个支持依赖感知形式化系统的评估基准。实验表明,尽管数据质量高,但任务极具挑战:强基线在定理陈述上的正确率最高仅9.8%,定义为16.7%。尤其在依赖深度最大的子集MA-Hard(700个实体)上,最佳模型正确率仅为2.6%。我们已将MathAtlas开源,供社区用于大尺度研究生数学自动形式化研究。

原文摘要 · Abstract (English)

Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains underexplored. In this paper, we introduce MathAtlas, the first large-scale autoformalization benchmark of in the wild graduate-level mathematics, containing ~52k theorems, definitions, exercises, examples, and proofs extracted from 103 graduate mathematics textbooks. MathAtlas is enriched with a mathematical dependency graph containing ~178k relations, and is the first autoformalization benchmark to include such relations, facilitating evaluation and development of dependency-aware autoformalization systems. Our extensive experiments show that MathAtlas is high quality but extremely challenging: strong baselines achieve at most 9.8% correctness on theorem statements and 16.7% on definitions. Furthermore, we find performance of state-of-the-art models degrades substantially with dependency depth: on MA-Hard, a subset of 700 entities with the deepest dependency trees, the best model achieves only 2.6% correctness for autoformalization on this challenging dataset. We release MathAtlas to the community as a benchmark set for large-scale autoformalization of graduate-level mathematics in the wild.

自动形式化数学AI依赖图谱研究生数学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。