优化大模型训练中的网络路由,提升通信效率
Routing for Large ML Models

- 通过量化全局通信效率,动态优化数据流路由
- 针对大模型训练的规律性通信模式,实现高效路径规划
- 适合大规模分布式训练场景,提升训练速度
训练大型语言模型(LLMs)及其他大规模机器学习模型,需要在数据中心网络中反复传输大量数据。训练过程引发的通信模式具有高度规律性和持久性,为优化数据流在网络中的路由方式提供了重要机会。本文提出一个算法框架,用于量化大规模模型训练背景下的全局网络效率,并据此周期性地优化路由策略。该方法旨在以全局性能指标为导向,提升数据传输效率,降低通信延迟,从而加速模型训练。
原文摘要 · Abstract (English)
Training large language models (LLMs), and other large machine learning models, involves repeated communication of large volumes of data across a data center network. The communication patterns induced by these training process exhibit high regularity and persistence, giving rise to significant opportunities for optimizing the manner in which flows are routed across the network. We present an algorithmic framework for \textit{quantifying} network-wide efficiency in the context of training LLMs (and other large-scale ML models), and for periodically \textit{optimizing} routing with respect to this global metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。