测试大模型在图论研究中的数学推理能力,发现只有顶尖模型能完成研究生级证明。
GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory
- 按难度分层设计63道图论题,覆盖本科到研究生阶段
- GPT-5在本科题上达95.8%准确率,研究生证明正确率达82%
- 揭示模型在复杂证明中存在推理不完整和评估分歧问题
大型语言模型(LLMs)在技术学科中日益用作自主学习助手,但其作为数学推理助手的可靠性仍不明确。我们提出GTBench,一个基于课程体系的图论数学研究助理评估基准,包含63道题目,分为三个难度递增组别:本科定义与基础性质(第1组)、算法追踪与结构推理(第2组)、研究生级证明构建(第3组)。题目源自经验证的学术资料,如Diestel的《图论》。我们对五款前沿模型——GPT-5、Claude Sonnet 4.6、Gemini 2.5 Flash-Lite、Llama 3.3 70B、Mistral Large 3——在零样本与思维链提示下进行评估,第1、2组采用精确匹配与大模型判卷,第3组采用人工专家与大模型联合评判。结果表明性能呈现显著层级:GPT-5在第1组接近上限(零样本95.8%),并在研究生证明中保持82%准确率;其余模型随难度大幅下降,其中Llama在第3组零样本人类评估中得分为0%。错误分析显示,第1、2组以“正确算法但执行错误”为主,第3组则出现推理不完整问题,并暴露出人类评审者与自动评判间系统性分歧,尤其在冗长或近完整的证明中(人对人一致性kappa=0.48–0.83)。GTBench为大模型图论推理提供了首个基于课程的评估框架,对人工智能在数学教育与科研治理具有直接意义。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning assistants remains poorly understood. We introduce GTBench, a curriculum-grounded benchmark for evaluating LLMs as mathematical research assistants in graph theory, comprising 63 problems organized into three groups of increasing difficulty: undergraduate definitions and basic properties (Group 1), algorithm tracing and structural reasoning (Group 2), and graduate-level proof construction (Group 3). Problems are sourced from verified academic materials including Diestel's Graph Theory. We evaluate five frontier models -- GPT-5, Claude Sonnet 4.6, Gemini 2.5 Flash-Lite, Llama 3.3 70B, and Mistral Large 3 -- under zero-shot and chain-of-thought prompting, using exact-match and LLM-as-judge evaluation for Groups 1 and 2, and a hybrid human expert and LLM-as-judge protocol for Group 3. Our results reveal a pronounced performance hierarchy: GPT-5 approaches ceiling on Group 1 (95.8% zero-shot) and maintains meaningful accuracy on graduate proofs (82%), while all other models degrade substantially with difficulty, with Llama achieving 0% under human evaluation on Group 3 zero-shot. Failure mode analysis shows that correct algorithm, wrong execution errors dominate Groups 1 and 2, while Group 3 additionally surfaces incomplete reasoning failures and reveals systematic disagreement between human evaluators and the automated judge, particularly on verbose or near-complete proofs (kappa = 0.48-0.83 across human pairs). GTBench provides the first curriculum-grounded evaluation framework for graph-theoretic reasoning in LLMs, with direct implications for the governance of AI tools in mathematical education and scientific research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。