arXiv:2510.01241cs.CL2025-10

构建双基准测试集,评估大模型在多层级数学推理中的真实能力

SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation

  • 设计结构感知的诊断题集与竞赛风格题集,含详细题目元数据
  • 最强模型在竞赛题上仅达44%,博士级难度准确率下降至高中水平的79%
  • 提供可复现的评测基准,适合评估数学推理模型的真实性能

大型语言模型在众多公开数学测评中表现强劲,但数学前沿能力的区分度正面临天花板效应。本文提出两个互补的评测基准:SKYLENAGE-ReasoningMATH,一个包含100道题、带长度、数值密度和符号复杂度等元数据的结构感知诊断集;以及SKYLENAGE-MATH,一个涵盖高中到博士四个阶段、共150道题的竞赛式题集,按七大学科分类。我们在统一设置下评估了15种主流LLM变体,分析了学科×模型和年级×模型的表现。在竞赛题集上,最强模型准确率为44%,次优为37%;准确率从高中到博士逐步下降,顶尖系统在博士与高中题间的保留率约为79%。在推理题集上,最佳模型整体达到81%准确率,最难子集揭示领先者与中游模型间存在明显鲁棒性差距。综上,我们发布SKYLENAGE-ReasoningMATH并报告SKYLENAGE-MATH的聚合结果;两者结合提供了具有校准难度、聚焦推理、覆盖广泛的数学评测基准,可作为未来数学推理评估的参考标准。

原文摘要 · Abstract (English)

Large language models (LLMs) now perform strongly on many public math suites, yet frontier separation within mathematics increasingly suffers from ceiling effects. We present two complementary benchmarks: SKYLENAGE-ReasoningMATH, a 100-item, structure-aware diagnostic set with per-item metadata on length, numeric density, and symbolic complexity; and SKYLENAGE-MATH, a 150-item contest-style suite spanning four stages from high school to doctoral under a seven-subject taxonomy. We evaluate fifteen contemporary LLM variants under a single setup and analyze subject x model and grade x model performance. On the contest suite, the strongest model reaches 44% while the runner-up reaches 37%; accuracy declines from high school to doctoral, and top systems exhibit a doctoral-to-high-school retention near 79%. On the reasoning set, the best model attains 81% overall, and hardest-slice results reveal clear robustness gaps between leaders and the mid-tier. In summary, we release SKYLENAGE-ReasoningMATH and report aggregate results for SKYLENAGE-MATH; together, SKYLENAGE provides a hard, reasoning-centered and broadly covering math benchmark with calibrated difficulty and rich metadata, serving as a reference benchmark for future evaluations of mathematical reasoning.

数学推理评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。