arXiv:2605.09063cs.CL2026-05被引 4

439道数学研究级题目,检验大模型真实科研推理能力。

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

论文配图:Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
图 1 · 摘自论文原文
  • 64位数学家原创439题,分挑战与拒答两子集。
  • 顶尖模型最高仅30.4%正确率,开放模型普遍低于15%。
  • 新增拒答测试,逼模型识别无解问题,填补评估空白。

在前沿大模型于国际数学奥林匹克竞赛中取得金牌级表现后,社区亟需更具挑战性的评估目标。相较于仅考察步骤推理的奥数题,研究级数学问题要求模型运用推理推动数学知识边界,成为更优替代。然而现有研究级基准稀缺——如Riemann Bench和FrontierMath-Tier 4分别仅有25和50题。为此,我们引入由64位数学家从零创作的Soohak基准,共439题,包含挑战子集与拒答子集。在挑战子集上,Gemini-3-Pro、GPT-5和Claude-Opus-4.5正确率分别为30.4%、26.4%和10.4%,领先开源模型(如Qwen3-235B、GPT-OSS-120B、Kimi-2.5)均未超15%。拒答子集则考察模型识别无解问题并主动拒绝的能力,当前所有模型最高未达50%,凸显该能力缺失。为防污染,数据集将于2026年底公开,评测可提前申请。

原文摘要 · Abstract (English)

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.

数学推理大模型评估研究级基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。