arXiv:2512.13978cs.AI2025-12被引 1

评测顶尖大模型在随机算法理论书上的证明能力,发现表现差异显著。

Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms

  • 用经典教材题目测试4个前沿模型的数学证明能力。
  • 顶级模型准确率达66%,其他模型仅40%左右。
  • 适合研究AI数学推理或教育辅助系统的人参考。

大语言模型在自动化数学推理和科学发现方面取得显著进展。尽管如此,对这些模型在研究生级数学理论基准上的严谨评估仍显不足。本文构建了一个综合性基准,测试GPT-5-Thinking、Gemini-3-Pro、Claude-Sonnet-4.5-Thinking和Grok-4四款前沿模型在《随机算法》(Motwani and Raghavan, [MR95])教材中的表现。要求模型生成一系列引理与习题的正式LaTeX证明。结果表明,顶级模型(Gemini和Claude)准确率约66%,展现出对概率方法和形式逻辑的良好掌握;其他模型一致性较差,约40%。我们还分析了生成证明的简洁性、幻觉率和逻辑结构差异。结果显示,尽管前沿模型已具备胜任研究生教学辅助与形式化的能力,但在严谨数学推导中的可靠性仍存在显著差异。代码与全部生成结果已开源,详见https://github.com/magiclinux/math_benchmark_probability。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has led to significant breakthroughs in automated mathematical reasoning and scientific discovery. Georgiev, G${ó}$mez-Serrano, Tao, and Wagner [GGSTW+25] demonstrate that AI systems can explore new constructions and improve existing bounds, illustrating the growing potential of LLMs to accelerate mathematical discovery. Similarly, Bubeck et al. [BCE+25] show that GPT-5 can meaningfully contribute to scientific workflows, from proposing hypotheses to generating proofs and analyses. Despite these advances, a rigorous evaluation of these models on canonical, graduate-level mathematical theory remains necessary to understand their baseline reasoning capabilities. In this paper, we present a comprehensive benchmark of four frontier models: GPT-5-Thinking, Gemini-3-Pro, Claude-Sonnet-4.5-Thinking, and Grok-4 against the classic curriculum of Randomized Algorithms by Motwani and Raghavan [MR95]. We tasked each model with generating formal LaTeX proofs for a series of lemmas and exercises spanning the textbook. We find that while the top-tier models (Gemini, and Claude) achieve a high accuracy rate (approx. 66%), demonstrating a robust grasp of probabilistic method and formal logic, other models lag significantly in consistency (approx. 40%). We provide a qualitative analysis of the generated proofs, highlighting differences in conciseness, hallucination rates, and logical structure. Our results suggest that while frontier models have reached a threshold of proficiency suitable for graduate-level pedagogical assistance and formalization, significant variance exists in their reliability for rigorous mathematical derivation. The code and the full set of LLM-generated responses are open-sourced and publicly available at https://github.com/magiclinux/math_benchmark_probability.

数学推理大模型评测随机算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。