arXiv:2504.09307cs.DCcs.AI2025-04中稿 · MLSys 2025被引 14

Lumos可精准预测大模型训练性能,助力高效配置优化

Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training

  • 基于实际运行轨迹建模,捕捉大模型训练行为
  • 在512张H100 GPU上平均误差仅3.3%
  • 支持新配置快速估测,适合系统优化与部署研究

在分布式环境中训练大语言模型面临模型执行、部署系统和可配置策略空间庞大的挑战。尽管存在多种优化技术,实际效率提升仍困难重重。准确的性能模型对指导优化和系统研究至关重要。我们提出Lumos,一个面向大规模大语言模型训练的迹驱动性能建模与估算工具包,能精确捕捉并预测现代大模型的执行行为。我们在包含最多512张NVIDIA H100 GPU的生产级ML集群上,使用多种GPT-3变体评估Lumos,结果表明其在不同模型与配置下可实现平均3.3%的执行时间重放误差,并准确还原其他运行时细节。此外,我们验证了其利用已有轨迹估算新配置性能的能力,显著提升模型与部署配置探索效率。

原文摘要 · Abstract (English)

Training LLMs in distributed environments presents significant challenges due to the complexity of model execution, deployment systems, and the vast space of configurable strategies. Although various optimization techniques exist, achieving high efficiency in practice remains difficult. Accurate performance models that effectively characterize and predict a model's behavior are essential for guiding optimization efforts and system-level studies. We propose Lumos, a trace-driven performance modeling and estimation toolkit for large-scale LLM training, designed to accurately capture and predict the execution behaviors of modern LLMs. We evaluate Lumos on a production ML cluster with up to 512 NVIDIA H100 GPUs using various GPT-3 variants, demonstrating that it can replay execution time with an average error of just 3.3%, along with other runtime details, across different models and configurations. Additionally, we validate its ability to estimate performance for new setups from existing traces, facilitating efficient exploration of model and deployment configurations.

大模型训练性能建模分布式系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。