arXiv:2508.06524cs.CLcs.AI2025-08被引 2

提出碳足迹预测框架,精准估算大模型训练的碳排放。

CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models

  • 基于硬件感知建模,融合分布式训练与能耗因素。
  • 在万亿参数规模下,实体碳排占比显著上升。
  • 适合关注大模型可持续性的研究者与工程师使用。

大型语言模型(LLMs)遵循神经缩放定律,性能提升依赖于快速扩张的计算预算,引发对前沿规模训练可持续性的担忧。现有碳排放估算方法主要依赖历史运行的回归分析,难以捕捉硬件异构性、分布式并行、通信开销和架构稀疏性等关键系统因素。我们提出 extit{CarbonScaling},一个面向前沿大模型训练的碳足迹缩放行为分析框架。该框架整合神经缩放定律、分布式训练策略、加速器与互连网络建模,以及运营碳与隐含碳核算,可估算可行的硬件配置及其对应的碳排放。CarbonScaling 同时建模张量、流水线、数据与专家并行,纳入内存、带宽、利用率与运行时间约束。实验验证表明,其精度显著优于基于回归的基线方法,并揭示在万亿参数规模下,隐含碳排放的重要性日益凸显。源代码: url{https://github.com/UnchartedRLab/CarbonScaling}。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly follow neural scaling laws that tie performance gains to rapidly expanding computational budgets, raising concerns about the sustainability of frontier-scale training. Existing carbon-estimation methods largely depend on regression over historical runs and fail to capture critical system-level factors, including hardware heterogeneity, distributed parallelism, communication overhead, and architectural sparsity. We present \textit{CarbonScaling}, a hardware-aware analytical framework for modeling the carbon scaling behavior of frontier LLM training. The framework integrates neural scaling laws, distributed training strategies, accelerator and interconnect modeling, and operational and embodied carbon accounting to estimate feasible hardware configurations and associated emissions. CarbonScaling jointly models tensor, pipeline, data, and expert parallelism while incorporating memory, bandwidth, utilization, and runtime constraints. Experimental validation demonstrates substantially higher fidelity than regression-based baselines and highlights the growing importance of embodied carbon at trillion-parameter scales. Source code: \url{https://github.com/UnchartedRLab/CarbonScaling}.

大模型碳足迹可持续性分布式训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。