测试AI在真正数学研究中的推理能力,发现当前模型表现远低于人类水平。
Riemann-Bench: A Benchmark for Moonshot Mathematics
- 由顶尖数学家设计难题,需数周独立求解
- 所有前沿模型平均得分不足10%,差距显著
- 私有基准防记忆泄露,真实评估研究能力
近期人工智能系统已在国际数学奥林匹克竞赛中达到金牌水平,展现出在竞赛类问题求解上的卓越能力。然而,竞赛数学仅涵盖数学推理的狭小范围:题目来自有限领域,依赖少量高级工具,常奖励巧妙技巧而非深层理论知识。我们提出Riemann-Bench,一个由常春藤盟校教授、研究生及奥数金牌得主共同设计的私有基准,用于评估AI在远超奥数边界的研究级数学推理能力。题目由作者独立花费数周求解,每题经两位领域专家双盲验证并提交唯一闭式解,由程序化验证器判定正确性。我们以无偏统计估计器,在每题100次独立运行下,评估前沿模型作为无约束研究代理的表现(支持编码、搜索与开放推理)。结果表明,所有模型得分均低于10%,暴露出竞赛解题与真实科研推理之间的巨大鸿沟。通过保持基准完全私有,确保性能衡量反映真实的数学能力而非训练数据记忆。
原文摘要 · Abstract (English)
Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving. However, competition mathematics represents only a narrow slice of mathematical reasoning: problems are drawn from limited domains, require minimal advanced machinery, and can often reward insightful tricks over deep theoretical knowledge. We introduce Riemann-Bench, a private benchmark of expert-curated problems designed to evaluate AI systems on research-level mathematics that goes far beyond the olympiad frontier. Problems are authored by Ivy League mathematics professors, graduate students, and PhD-holding IMO medalists, and routinely took their authors weeks to solve independently. Each problem undergoes double-blind verification by two independent domain experts who must solve the problem from scratch, and yields a unique, closed-form solution assessed by programmatic verifiers. We evaluate frontier models as unconstrained research agents, with full access to coding tools, search, and open-ended reasoning, using an unbiased statistical estimator computed over 100 independent runs per problem. Our results reveal that all frontier models currently score below 10%, exposing a substantial gap between olympiad-level problem solving and genuine research-level mathematical reasoning. By keeping the benchmark fully private, we ensure that measured performance reflects authentic mathematical capability rather than memorization of training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。