arXiv:2602.15327cs.LGcs.AI2026-02

用大规模实验揭示大模型能力随算力演进的规律,帮开发者预判性能上限。

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

  • 基于2022-2026年数千个模型检查点,建立算力与性能的量化预测模型。
  • 在10^24 FLOPs预算下,IFEval准确率可达83%,MATH Lvl 5达54%。
  • 提出高效采样算法,仅需5%-20%评估成本即可逼近完整性能边界。

机器学习模型性能提升常源于竞争与应用。为部署提供可指导的缩放定律:给定预训练算力预算,当前后训练方法下可实现的下游准确率是多少?该映射关系随领域演进是否稳定?我们通过涵盖2022-2026年、6个基准测试的5000个现有与2000个新评估模型检查点的大规模观测实验,采用平滑分位数回归与单调饱和型S形参数化,估计了以对数预训练FLOPs为变量的能力边界(高条件分位数)。通过早期模型拟合、后期模型验证,四类任务外推误差低于2%,数学推理能力则持续提升。例如,在10^24 FLOPs预算下,预计可达到IFEval准确率0.83,MATH Lvl 5为0.54。进一步分析任务依赖性饱和现象及数学推理中的数据污染影响。最后,提出一种平衡的I-最优采样算法,仅需约20%(部分任务低至5%)的参数量加权评估预算,即可恢复接近全数据的性能前沿,且校准效果相当。本工作发布Proteus-2k——最新模型性能评估数据集,并提供将算力预算转化为可靠性能预期的实用方法,以及监测能力边界随时间变化的工具。

原文摘要 · Abstract (English)

Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scaling laws: given a pre-training compute budget, what downstream accuracy is attainable with contemporary post-training practice, and how stable is that mapping as the field evolves? Using large-scale observational evaluations with 5k existing and 2k newly evaluated model checkpoints spanning 2022-2026 across six benchmarks, we estimate capability boundaries, high conditional quantiles of benchmark scores as a function of log pre-training FLOPs, via smoothed quantile regression with a monotone, saturating sigmoid parameterization. We validate temporal reliability by fitting on earlier model generations and evaluating on later releases: across four of six tasks, the out-of-distribution coverage error remains below 2%, while math reasoning exhibits a consistently advancing boundary over time. For instance, at a budget of 10^24 FLOPs, the estimated attainable accuracies are 0.83 on IFEval and 0.54 on MATH Lvl 5. We then extend our approach to analyze task-dependent saturation and to probe contamination-related shifts on math reasoning tasks. Finally, we introduce a balanced I-optimal sampling algorithm that recovers near-full-data frontiers using roughly 20% of the parameter-count-weighted evaluation budget, as low as 5% on some tasks, while maintaining comparable calibration. Together, our work releases Proteus-2k, the latest model performance evaluation dataset, and introduces a practical methodology for translating compute budgets into reliable performance expectations and for monitoring when capability boundaries shift across time.

大模型缩放定律性能预测评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。