arXiv:2512.00193cs.AI2025-12被引 23

建立跨基准统一量化体系,实现对AI能力的长期追踪与对比。

A Rosetta Stone for AI Benchmarks

  • 构建统一数值尺度,让不同时间、不同任务的模型可比。
  • 发现算法效率提升速度高于以往估计,但整体趋势一致。
  • 能检测出AI进步中的加速现象,适合研究者跟踪技术演进。

大多数AI基准在发布后数月或数年内即趋于饱和,难以追踪AI能力的长期趋势。为此,我们构建了一个统计框架,将不同基准拼接起来,使模型能力与基准难度处于同一数值尺度上。该框架如同‘罗塞塔石碑’,可在不依赖训练计算量或能力随时间演变假设的前提下,跨时间、跨任务比较模型表现。我们展示了三个应用:首先,测量AI进步速度并预测未来能力;其次,估算算法效率提升速率,结果高于以往估计但总体一致;最后,证明该方法可检测到AI进展的快速加速现象。

原文摘要 · Abstract (English)

Most AI benchmarks saturate within years or even months after they are introduced, making it hard to study long-run trends in AI capabilities. To address this challenge, we build a statistical framework that stitches benchmarks together, putting model capabilities and benchmark difficulties on a single numerical scale. This acts as a "Rosetta Stone", allowing us to compare models across a wide range of abilities and time, even if they are not evaluated on the same benchmarks. Moreover, this works without assuming how capabilities evolve across time or with training compute. We demonstrate three applications of this framework. First, we use it to measure the speed of AI progress over time, and to forecast future AI capabilities. Second, we estimate the rate of improvements in algorithmic efficiency, finding estimates that are higher, but broadly consistent with prior work. Finally, we find that our approach can be used to detect rapid accelerations in AI progress.

AI评估长期趋势基准统一

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。