arXiv:2508.10173cs.LGcs.CY2025-08

用重要基准训练可提升大模型推理能力,而非仅靠模型变大或算法优化。

Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1

  • 以关键基准作为学习课程,指导模型训练
  • 在Humanity's Last Exam任务中,表现提升源于基准选择而非模型尺寸
  • 部分基准实为训练课程,适合关注通用推理的开发者

推理型语言模型的评估日益重要,因其能在任务完成前生成新颖的中间步骤,并可能实现更好泛化。随着推理成为大模型的下一个规模维度,对关键任务能力的深入研究必不可少。我们发现,性能提升不仅来自测试时的算法改进或模型规模,更源于使用有影响力的基准作为学习课程。我们称之为‘基准驱动的AI选择’,并在DeepSeek-R1上通过Humanity's Last Exam的顺序决策问题验证了其效果。以关键基准引导发展,将评估转化为学习,使测试任务的新颖性成为衡量推理模型泛化能力的关键。因此,某些基准可视为训练课程,而非纯粹的测试集。

原文摘要 · Abstract (English)

Evaluation of reasoning language models gained importance after it was observed that they can combine their existing capabilities into novel traces of intermediate steps before task completion and that the traces can sometimes help them to generalize better than past models. As reasoning becomes the next scaling dimension of large language models, careful study of their capabilities in critical tasks is needed. We show that better performance is not always caused by test-time algorithmic improvements or model sizes but also by using impactful benchmarks as curricula for learning. We call this benchmark-driven selection of AI and show its effects on DeepSeek-R1 using our sequential decision-making problem from Humanity's Last Exam. Steering development of AI by impactful benchmarks trades evaluation for learning and makes novelty of test tasks key for measuring generalization capabilities of reasoning models. Consequently, some benchmarks could be seen as curricula for training rather than unseen test sets.

推理模型基准评估训练课程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。