arXiv:2507.07931cs.AIcs.CY2025-07被引 2

计算资源越少的模型,未来越可能追上顶尖模型表现。

Meek Models Shall Inherit the Earth

  • 在固定分布预测任务下,增加算力带来的性能提升逐渐减少。
  • 当前算力扩展模式下,大公司优势将随时间减弱至几乎消失。
  • 适合关注模型公平性、政策制定与小型团队技术布局的人阅读。

过去十年,少数公司通过大规模扩展AI系统,导致模型性能不平等。本文指出,与普遍认知相反,算力扩展的边际收益递减将推动各类模型能力趋于收敛。即算力有限的‘弱模型’将逐步接近最优模型的表现。我们构建了一个模型,证明在固定分布的下一个词预测目标下,额外算力带来的能力提升显著衰减。基于当前扩展实践,即使某些公司能以指数级速度扩展模型,其能力优势也将逐渐消失。本文还提供证据表明,训练损失等代理指标可有效反映实际能力,并分析了历史模型性能差异数据。最后,鉴于弱模型能力的提升,我们呼吁重新审视人工智能战略与政策,明确这一转变的影响领域。

原文摘要 · Abstract (English)

The past decade has seen incredible scaling of AI systems by a few companies, leading to inequality in AI model performance. This paper argues that, contrary to prevailing intuition, the diminishing returns to compute scaling will lead to a convergence of AI model capabilities. In other words, meek models (those with limited computation budget) shall inherit the earth, approaching the performance level of the best models overall. We develop a model illustrating that under a fixed-distribution next-token objective, the marginal capability returns to raw compute shrink substantially. Given current scaling practices, we argue that these diminishing returns are strong enough that even companies that can scale their models exponentially faster than other organizations will eventually have little advantage in capabilities. As part of our argument, we give several reasons that proxies like training loss differences capture important capability measures using evidence from benchmark data and theoretical performance models. In addition, we analyze empirical data on the capability difference of AI models over time. Finally, in light of the increasing ability of meek models, we argue that AI strategy and policy require reexamination, and we outline the areas this shift will affect.

模型收敛算力效率能力公平战略评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。