arXiv:2509.24827cs.LGcs.AI2025-09被引 1

测试大模型解数学竞赛题能力,发现顶级模型表现优异但仍有局限。

Putnam-like dataset summary: LLMs as mathematical competition contestants

  • 用96道类普特南竞赛题评估大模型数学推理能力
  • 顶级模型得分接近人类水平,但2024年真题表现下降
  • 模型常出现两极化得分,难以给出严谨证明

本文总结了谷歌深度求索发布的类普特南竞赛数据集结果。该数据集包含96道具有普特南竞赛风格的原创题目及576个由大语言模型生成的解法。我们分析了模型在这些题目上的表现,以验证其解决数学竞赛问题的能力。结果显示,顶尖模型(尤其是Gemini 2.5 Pro)得分较高,展现出强大的数学推理能力,但在2024年普特南竞赛题目上表现有所下降。分析揭示了模型行为的显著特征,包括得分分布呈双峰形态,以及在提供完全严谨的数学推导方面存在困难。

原文摘要 · Abstract (English)

In this paper we summarize the results of the Putnam-like benchmark published by Google DeepMind. This dataset consists of 96 original problems in the spirit of the Putnam Competition and 576 solutions generated by LLMs. We analyze the performance of models on this set of problems to verify their ability to solve problems from mathematical contests. We find that top models, particularly Gemini 2.5 Pro, achieve high scores, demonstrating strong mathematical reasoning capabilities, although their performance was lower on problems from the 2024 Putnam competition. The analysis highlights distinct behavioral patterns among models, including bimodal scoring distributions and challenges in providing fully rigorous justifications.

数学推理大模型评测竞赛题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。