研究大模型算法创新的算力需求,发现算力限制难阻创新步伐。
Compute Requirements for Algorithmic Innovation in Frontier AI Models
- 统计36项大模型预训练算法创新的算力消耗
- 算力需求每年翻倍,但算力上限仍可支撑半数创新
- 适合关注AI发展瓶颈与算力政策的研究者
大型语言模型预训练中的算法创新已大幅降低达到特定能力水平所需的总计算量。本文通过实证研究算法创新的算力需求,梳理了Llama 3和DeepSeek-V3中使用的36项预训练算法创新,并估算每项创新在开发过程中所用的总浮点运算量(FLOP)及硬件算力(FLOP/s)。结果显示,高资源投入的创新其算力需求每年翻倍。我们进一步利用该数据集分析算力上限对创新的影响,结果表明,仅靠算力限制难以显著延缓人工智能算法进步。即使设定严格算力上限——如将总操作量限制在GPT-2的训练算力,或硬件容量限制在8块H100 GPU——仍足以支持所列创新的一半。
原文摘要 · Abstract (English)
Algorithmic innovation in the pretraining of large language models has driven a massive reduction in the total compute required to reach a given level of capability. In this paper we empirically investigate the compute requirements for developing algorithmic innovations. We catalog 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3. For each innovation we estimate both the total FLOP used in development and the FLOP/s of the hardware utilized. Innovations using significant resources double in their requirements each year. We then use this dataset to investigate the effect of compute caps on innovation. Our analysis suggests that compute caps alone are unlikely to dramatically slow AI algorithmic progress. Even stringent compute caps -- such as capping total operations to the compute used to train GPT-2 or capping hardware capacity to 8 H100 GPUs -- could still have allowed for half of the cataloged innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。