构建首个中印奥数题人工验证形式化基准,测试大模型数学推理能力。
IndiMathBench: Autoformalizing Mathematical Reasoning Problems with a Human Touch
- 用AI辅助+专家验证流程,将312道印度奥数题转为Lean 4形式化定理。
- 大模型形式化准确率低,语义正确性远低于语法正确性。
- 适合评估数学推理与形式化能力,推动自动化证明发展。
即使在大型语言模型时代,可靠的形式化依然充满挑战,主要瓶颈在于高质量训练数据稀缺。专家标注需大量时间及数学与定理证明双重专长。本文提出IndiMathBench,一个经人工验证的基准,用于评估数学定理证明能力,通过AI赋能的人机协作流程,在Lean中形式化自然语言问题。该基准包含312个配对的正式Lean 4定理及其对应的非形式化题干,均来自印度数学奥林匹克竞赛。利用基于类别检索、迭代编译器反馈和多模型集成的管道,生成候选形式化结果,由专家通过交互式仪表板高效验证,并辅以自动化质量摘要。对多个前沿模型的评估显示,形式化仍具挑战,语法正确性与语义正确性间存在显著差距,即使经过迭代优化,定理证明成功率依然偏低,表明该基准构成了一项极具挑战性的数学推理测试平台。IndiMathBench开源地址:https://github.com/prmbiy/IndiMathBench。
原文摘要 · Abstract (English)
Reliable autoformalization remains challenging even in the era of large language models (LLMs). The scarcity of high-quality training data is a major bottleneck. Expert annotation requires substantial time and deep expertise in both mathematics and theorem proving. We introduce IndiMathBench, a human-verified benchmark designed to evaluate mathematical theorem proving, curated using an AI-powered human-assisted pipeline for formalizing natural language problems in Lean. IndiMathBench is composed of 312 formal Lean 4 theorems paired with their corresponding informal problem statements, sourced from Indian Mathematics Olympiads. Through category-based retrieval, iterative compiler feedback, and multi-model ensembles, our pipeline generates candidate formalizations that experts efficiently validate via an interactive dashboard with automated quality summaries. Evaluation across multiple frontier models demonstrates that autoformalization remains challenging, with substantial gaps between syntactic validity and semantic correctness, while theorem proving success rates remain low even with iterative refinement, demonstrating that \benchmark~presents a challenging testbed for mathematical reasoning. IndiMathBench is available at https://github.com/prmbiy/IndiMathBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。