Kareus同时优化大模型训练的动态与静态能耗,显著降低能源消耗或加速训练。
Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
- 通过细粒度核调度与频率调节协同优化能耗
- 相同训练时间下节能最高达28.3%,相同能耗下提速最高达27.5%
- 适用于需要高效能比的大模型训练场景
AI计算需求以空前速度增长,但能源供给未能同步,导致能源成为昂贵且需显式管理的资源。尽管近期研究在大模型训练优化上取得进展,但多仅关注动态或静态能耗之一。我们发现,细粒度核调度与频率调节会联合且相互依赖地影响两类能耗。基于此,我们设计了Kareus——一个通过联合优化动态与静态能耗来突破时-能权衡边界的训练系统。Kareus将难以求解的联合优化问题分解为局部、基于分区的子问题,并采用多轮多目标优化算法寻找最优执行调度,可使训练能耗降低最多28.3%(同训练时间),或训练时间缩短最多27.5%(同能耗)。
原文摘要 · Abstract (English)
The computing demand of AI is growing at an unprecedented rate, but energy supply is not keeping pace. As a result, energy has become an expensive and contended resource that requires explicit management and optimization. Although recent works have made significant progress in large model training optimization, they focus on optimizing either dynamic or static energy consumption. We find that fine-grained kernel scheduling and frequency scaling jointly and interdependently impact both dynamic and static energy consumption. Based on this finding, we design Kareus, a training system that pushes the time-energy tradeoff frontier by optimizing both aspects. Kareus decomposes the intractable joint optimization problem into local, partition-based subproblems. It then uses a multi-pass multi-objective optimization algorithm to find execution schedules that push the time-energy tradeoff frontier. Compared to the state of the art, Kareus reduces training energy by up to 28.3% at the same training time, or reduces training time by up to 27.5% at the same energy consumption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。