量化AI训练的材料消耗,揭示大模型背后的资源代价
From FLOPs to Footprints: The Resource Cost of Artificial Intelligence
- 通过分析GPU元素组成与计算负载,建立算力与材料消耗的关联
- 训练GPT-4需1174至8800块A100 GPU,最多产生7吨有毒元素
- 提升算力利用率和延长硬件寿命可减少93%的材料需求
随着计算需求增长,评估AI环境足迹需超越能耗与水耗,纳入专用硬件的材料需求。本研究通过将计算工作量与物理硬件需求关联,量化了AI训练的材料足迹。采用电感耦合等离子体发射光谱法分析Nvidia A100 SXM 40 GB GPU的元素组成,识别出32种元素,其中约90%为重金属,铜、铁、锡、硅、镍占质量主导。通过多步方法,结合不同生命周期下每块GPU的计算吞吐量,模拟不同训练效率下特定AI模型的硬件需求。情景分析显示,训练GPT-4需1,174至8,800块A100 GPU,对应提取并最终处置最多7吨有毒元素。软硬件协同优化可显著降低材料消耗:将模型浮点运算利用率(MFU)从20%提升至60%,可减少67%的GPU需求;硬件寿命从1年延长至3年,效果相当;两者结合可使需求降低高达93%。研究指出,如GPT-3.5到GPT-4的性能提升,带来不成比例的材料成本。强调未来AI发展必须纳入资源效率与环境责任考量。
原文摘要 · Abstract (English)
As computational demands continue to rise, assessing the environmental footprint of AI requires moving beyond energy and water consumption to include the material demands of specialized hardware. This study quantifies the material footprint of AI training by linking computational workloads to physical hardware needs. The elemental composition of the Nvidia A100 SXM 40 GB graphics processing unit (GPU) was analyzed using inductively coupled plasma optical emission spectroscopy, which identified 32 elements. The results show that AI hardware consists of about 90% heavy metals and only trace amounts of precious metals. The elements copper, iron, tin, silicon, and nickel dominate the GPU composition by mass. In a multi-step methodology, we integrate these measurements with computational throughput per GPU across varying lifespans, accounting for the computational requirements of training specific AI models at different training efficiency regimes. Scenario-based analyses reveal that, depending on Model FLOPs Utilization (MFU) and hardware lifespan, training GPT-4 requires between 1,174 and 8,800 A100 GPUs, corresponding to the extraction and eventual disposal of up to 7 tons of toxic elements. Combined software and hardware optimization strategies can reduce material demands: increasing MFU from 20% to 60% lowers GPU requirements by 67%, while extending lifespan from 1 to 3 years yields comparable savings; implementing both measures together reduces GPU needs by up to 93%. Our findings highlight that incremental performance gains, such as those observed between GPT-3.5 and GPT-4, come at disproportionately high material costs. The study underscores the necessity of incorporating material resource considerations into discussions of AI scalability, emphasizing that future progress in AI must align with principles of resource efficiency and environmental responsibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。