测试多个AI系统解决10个真实数学研究问题的能力
First Proof Second Batch
- 用10个真实数学研究问题评估AI解题能力
- 涉及拓扑、数论、组合等多个数学领域
- 适合关注AI在数学研究中潜力的学者
为评估当前AI系统解决研究级数学问题的能力,我们对多个AI系统在十个涵盖广泛数学领域的研究问题上进行了测试;这些问题源自贡献者的研究过程。本文包含问题描述、测试方法及结果。我们提供了补充文档链接,包括人类解法、AI生成解法以及AI解法的评审报告和日志。这十个问题由以下数学家提供:(1) Dariusz Kalociński 和 Theodore A. Slaman,(2) Richard Schwartz,(3) Aleksa Milojevic 和 Benny Sudakov,(4) Larry Guth,(5) Oleg Butkovsky、Jonathan Mattingly 与 Lorenzo Zambotti,(6) Joshua Evan Greene 与 Duncan McCoy,(7) Sucharit Sarkar,(8) Sam Payne 与 Jidong (Jayden) Wang,(9) Sylvie Corteel 与 John Lentfer,(10) Srivatsav Kunnawalkam Elayavalli。
原文摘要 · Abstract (English)
To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set of ten problems in a broad range of mathematical fields; these problems arose naturally in the research process of the contributors. This document includes the problems, our methodology, and the results of our testing. We provide links to supplementary documents including the human solutions, the AI-generated solutions, and the referee reports and logs for the AI-generated solutions. The ten problems were contributed by the following mathematicians: (1) Dariusz Kalociński and Theodore A. Slaman, (2) Richard Schwartz, (3) Aleksa Milojevic and Benny Sudakov, (4) Larry Guth, (5) Oleg Butkovsky, Jonathan Mattingly, and Lorenzo Zambotti, (6) Joshua Evan Greene and Duncan McCoy, (7) Sucharit Sarkar, (8) Sam Payne and Jidong (Jayden) Wang, (9) Sylvie Corteel and John Lentfer, (10) Srivatsav Kunnawalkam Elayavalli.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。