100道高阶数学题测试大模型,仅剩2题未解。
Benchmarks in Leipzig
- 集结49位数学家构建100道有答案的高阶数学题集
- 三阶段评估后,大模型仅2题无法解答,成功率超98%
- 适合关注AI数学推理能力的研究者和开发者
2026年4月1日至5月15日,49名数学家共同构建了一个包含已知答案的研究级数学问题数据集。主要工作在德国莱比锡马克斯·普朗克数学科学研究所举行的为期三天的‘莱比锡基准研讨会’中完成,共有35人参与。本文呈现了最终收集的100道问题。我们分三个阶段对这些问题进行了评估:首先由五款最先进的大语言模型进行单次尝试;随后对其中三款模型进行每模型20次运行的重复测试;最后使用两款重型推理模型进行三次尝试。第一阶段后仍有41题未解;第二阶段后减少至16题;第三阶段结束时仅剩2题未解。这表明大语言模型的数学推理能力正在显著提升。
原文摘要 · Abstract (English)
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100 questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。