arXiv:2411.04872cs.AI2024-11被引 233

构建数学前沿难题基准,测试AI解题能力

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

  • 由专家设计数百道高难度数学题,覆盖现代数学主要领域
  • 当前顶尖AI模型仅能解决不足2%的问题,差距显著
  • 适合评估AI在数学推理上的进展,尤其面向专家级能力

我们提出FrontierMath,一个由专业数学家设计并验证的、包含数百道原创且极富挑战性的数学问题的基准测试集。题目涵盖现代数学主要分支,从数论与实分析中的计算密集型问题,到代数几何与范畴论中的抽象问题。典型题目需相关领域研究者投入数小时,高端题目甚至需数日才能解决。FrontierMath采用全新未发表题目,并结合自动化验证,有效避免数据泄露风险,确保评估可靠性。当前最先进AI模型仅能正确解答不足2%的问题,揭示出人工智能与数学界实际能力之间的巨大鸿沟。随着AI向专家级数学能力演进,FrontierMath提供了一个严谨的测试平台,用于量化其进展。

原文摘要 · Abstract (English)

We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most major branches of modern mathematics -- from computationally intensive problems in number theory and real analysis to abstract questions in algebraic geometry and category theory. Solving a typical problem requires multiple hours of effort from a researcher in the relevant branch of mathematics, and for the upper end questions, multiple days. FrontierMath uses new, unpublished problems and automated verification to reliably evaluate models while minimizing risk of data contamination. Current state-of-the-art AI models solve under 2% of problems, revealing a vast gap between AI capabilities and the prowess of the mathematical community. As AI systems advance toward expert-level mathematical abilities, FrontierMath offers a rigorous testbed that quantifies their progress.

数学推理基准测试AI评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。