用最新数学研究论文中的定理自动构建动态评测基准,评估大模型的科研级证明能力。
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
- 从arXiv论文自动提取定理,重写为自包含命题,确保前提与定义清晰。
- 当前顶级大模型在定理证明上准确率仅10%-15%(pass@1),距离人类水平差距大。
- 适合关注大模型数学推理、自动化证明研究的研究者使用。
我们提出一种新方法,用于评估大语言模型在科研级数学任务中的能力。现有评测主要依赖静态的手工筛选的竞赛或教材题目作为数学研究的代理指标。本文建立了一个可更新的基准,直接评估模型对最新数学研究成果的理解与证明能力。该基准通过自动化流程从arXiv中提取定理,并将其重写为包含所有前提和必要定义的自包含陈述。该方法可定期更新,引入来自真实数学研究的新问题,同时保留历史版本用于训练而不影响未来评估的公平性。我们在当前最先进的大模型上进行测试,结果显示其在定理证明任务上的准确率约为10-15%(pass@1),表明大模型在科研级数学证明方面仍存在巨大提升空间。
原文摘要 · Abstract (English)
We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of contest or textbook-style problems as proxies for mathematical research. Instead, we establish an updatable benchmark evaluating models directly on the latest research results in mathematics. This consists of an automatic pipeline that extracts lemmas from arXiv and rewrites them into self-contained statements by making all assumptions and required definitions explicit. It results in a benchmark that can be updated regularly with new problems taken directly from human mathematical research, while previous instances can be used for training without compromising future evaluations. We benchmark current state-of-the-art LLMs, which obtain around 10-15$\%$ accuracy in theorem proving (pass@1) depending on the model, showing that there is currently a large margin of progression for LLMs to reach human-level proving capabilities in a research context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。