arXiv:2511.21622cs.LGcs.AI2025-11被引 6

算法效率提升主要依赖计算规模,而非单一模型改进。

On the Origin of Algorithmic Progress in AI

  • 通过缩放实验发现算法效率随计算量增大而显著提升
  • 2012-2023年算法提升达6930倍,主要来自LSTM到Transformer的转变
  • 小模型效率增长远低于预期,评估需考虑计算规模参考

2012至2023年间,算法被估算使AI训练FLOP效率提升了22,000倍 [Ho et al., 2024]。我们对关键创新进行小规模消融实验,仅能解释不足10倍增益;文献调研显示未包含的其他创新也贡献不足10倍,总计不足100倍。为此开展缩放实验,发现大量效率差距源于具有规模依赖性优化的算法。特别是对比LSTM与Transformer,二者在计算最优缩放律中呈现指数差异,而其他多数创新则无明显缩放差异。结果表明:算法效率提升与计算规模密切相关,与常规假设相反。结合实验外推与文献估计,我们解释了6,930倍效率提升,其中从LSTM到Transformer的过渡贡献最大。研究揭示:小模型的算法进步远慢于以往认知,算法效率衡量高度依赖参考计算规模。

原文摘要 · Abstract (English)

Algorithms have been estimated to increase AI training FLOP efficiency by a factor of 22,000 between 2012 and 2023 [Ho et al., 2024]. Running small-scale ablation experiments on key innovations from this time period, we are able to account for less than 10x of these gains. Surveying the broader literature, we estimate that additional innovations not included in our ablations account for less than 10x, yielding a total under 100x. This leads us to conduct scaling experiments, which reveal that much of this efficiency gap can be explained by algorithms with scale-dependent efficiency improvements. In particular, we conduct scaling experiments between LSTMs and Transformers, finding exponent differences in their compute-optimal scaling law while finding little scaling difference for many other innovations. These experiments demonstrate that - contrary to standard assumptions - an algorithm's efficiency gains are tied to compute scale. Using experimental extrapolation and literature estimates, we account for 6,930x efficiency gains over the same time period, with the scale-dependent LSTM-to-Transformer transition accounting for the majority of gains. Our results indicate that algorithmic progress for small models has been far slower than previously assumed, and that measures of algorithmic efficiency are strongly reference-dependent.

算法效率缩放定律TransformerFLOP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。