arXiv:2604.06802cs.AI2026-04中稿 · ICLR被引 7

测试AI在真正数学研究中的推理能力,发现当前模型表现远低于人类水平。

Riemann-Bench: A Benchmark for Moonshot Mathematics

  • 由顶尖数学家设计难题,需数周独立求解
  • 所有前沿模型平均得分不足10%,差距显著
  • 私有基准防记忆泄露,真实评估研究能力

近期人工智能系统已在国际数学奥林匹克竞赛中达到金牌水平,展现出在竞赛类问题求解上的卓越能力。然而,竞赛数学仅涵盖数学推理的狭小范围:题目来自有限领域,依赖少量高级工具,常奖励巧妙技巧而非深层理论知识。我们提出Riemann-Bench,一个由常春藤盟校教授、研究生及奥数金牌得主共同设计的私有基准,用于评估AI在远超奥数边界的研究级数学推理能力。题目由作者独立花费数周求解,每题经两位领域专家双盲验证并提交唯一闭式解,由程序化验证器判定正确性。我们以无偏统计估计器,在每题100次独立运行下,评估前沿模型作为无约束研究代理的表现(支持编码、搜索与开放推理)。结果表明,所有模型得分均低于10%,暴露出竞赛解题与真实科研推理之间的巨大鸿沟。通过保持基准完全私有,确保性能衡量反映真实的数学能力而非训练数据记忆。

原文摘要 · Abstract (English)

Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving. However, competition mathematics represents only a narrow slice of mathematical reasoning: problems are drawn from limited domains, require minimal advanced machinery, and can often reward insightful tricks over deep theoretical knowledge. We introduce Riemann-Bench, a private benchmark of expert-curated problems designed to evaluate AI systems on research-level mathematics that goes far beyond the olympiad frontier. Problems are authored by Ivy League mathematics professors, graduate students, and PhD-holding IMO medalists, and routinely took their authors weeks to solve independently. Each problem undergoes double-blind verification by two independent domain experts who must solve the problem from scratch, and yields a unique, closed-form solution assessed by programmatic verifiers. We evaluate frontier models as unconstrained research agents, with full access to coding tools, search, and open-ended reasoning, using an unbiased statistical estimator computed over 100 independent runs per problem. Our results reveal that all frontier models currently score below 10%, exposing a substantial gap between olympiad-level problem solving and genuine research-level mathematical reasoning. By keeping the benchmark fully private, we ensure that measured performance reflects authentic mathematical capability rather than memorization of training data.

数学推理基准测试AI研究Riemann-Bench

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。