arXiv:2501.02156cs.LGcs.AI2025-01被引 3

提出效率竞争新框架,揭示模型扩展中算力与效率的博弈关系。

The Race to Efficiency: A New Perspective on AI Scaling Laws

  • 构建考虑时间和效率的相对损失方程,突破传统静态算力假设。
  • 若无持续效率提升,先进模型需数千年训练或超大规模GPU集群。
  • 只要效率每两年翻倍,即可维持近指数级进步,适合长期规划者参考。

随着大规模AI模型的发展,训练成本不断攀升,维持进步愈发困难。经典缩放定律(如Kaplan等, 2020;Hoffmann等, 2022)仅基于固定算力预算预测训练损失,忽略了时间与效率因素。本文提出相对损失方程,一个兼顾时间和效率的新型框架,拓展了传统缩放定律。模型表明,若无持续效率提升,先进性能可能需要数千年训练时间或不切实际的大规模GPU集群。然而,只要‘效率翻倍速率’与摩尔定律相当,近指数级进展仍可实现。通过形式化这一效率竞赛,我们为平衡前期巨额算力投入与全栈持续优化提供了量化路线图。实证趋势显示,持续效率提升可使AI缩放延续至未来十年,为理解经典缩放定律中的回报递减现象提供了新视角。

原文摘要 · Abstract (English)

As large-scale AI models expand, training becomes costlier and sustaining progress grows harder. Classical scaling laws (e.g., Kaplan et al. (2020), Hoffmann et al. (2022)) predict training loss from a static compute budget yet neglect time and efficiency, prompting the question: how can we balance ballooning GPU fleets with rapidly improving hardware and algorithms? We introduce the relative-loss equation, a time- and efficiency-aware framework that extends classical AI scaling laws. Our model shows that, without ongoing efficiency gains, advanced performance could demand millennia of training or unrealistically large GPU fleets. However, near-exponential progress remains achievable if the "efficiency-doubling rate" parallels Moore's Law. By formalizing this race to efficiency, we offer a quantitative roadmap for balancing front-loaded GPU investments with incremental improvements across the AI stack. Empirical trends suggest that sustained efficiency gains can push AI scaling well into the coming decade, providing a new perspective on the diminishing returns inherent in classical scaling.

AI缩放效率优化算力经济

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。