arXiv:2602.22760cs.DCcs.AI2026-02

利用风电光伏弃电时段训练大模型,省电又减碳。

Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study

  • 在弃电窗口期弹性调度多地GPU集群,实现分布式训练
  • 561M参数模型训练减排90%以上,质量与单点训练相当
  • 适合关注绿色AI、低碳计算的科研与工程团队

训练大语言模型(LLM)需要大量算力和能源。与此同时,可再生能源常产生超出电网消纳能力的电力,导致弃电现象。这些弃电时段恰好是机遇:若将训练任务对齐于弃电窗口,即可利用清洁且廉价的电力完成预训练。本技术报告提出一个系统,在区域弃电期间跨地理分布的GPU集群上执行全参数LLM训练,根据节点可用性动态切换本地单站点训练与联邦多站点同步。原型基于Flower框架,在三个集群上训练了561M参数的Transformer模型,弃电时段数据来自真实边际碳强度记录。初步结果表明,弃电感知调度在保持训练质量的同时,将运营排放降至单站点基线的5-12%。

原文摘要 · Abstract (English)

Training large language models (LLMs) requires substantial compute and energy. At the same time, renewable energy sources regularly produce more electricity than the grid can absorb, leading to curtailment, the deliberate reduction of clean generation that would otherwise go to waste. These periods represent an opportunity: if training is aligned with curtailment windows, LLMs can be pretrained using electricity that is both clean and cheap. This technical report presents a system that performs full-parameter LLM training across geo-distributed GPU clusters during regional curtailment windows, elastically switching between local single-site training and federated multi-site synchronization as sites become available or unavailable. Our prototype trains a 561M-parameter transformer model across three clusters using the Flower federated learning framework, with curtailment periods derived from real-world marginal carbon intensity traces. Preliminary results show that curtailment-aware scheduling preserves training quality while reducing operational emissions to 5-12% of single-site baselines.

绿色计算分布式训练低碳AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。