arXiv:2509.03263cs.LGcs.AI2025-09被引 1

分析GPU训练AI的效率瓶颈,发现优化配置可提升性能与能效平衡。

Estudio de la eficiencia en la escalabilidad de GPUs para el entrenamiento de Inteligencia Artificial

  • 基于MLPerf Training v4.1数据,对比BERT等四类模型的训练表现。
  • 发现存在性能与效率的平衡点,可缩短训练时间并提升资源利用率。
  • 适合关注AI训练硬件优化的研究者与工程团队参考。

大规模深度学习模型的训练已成为科学界和产业界的关键挑战。尽管大量使用GPU能显著加速训练过程,但这一做法对效率产生负面影响。本文基于MLPerf Training v4.1的四个工作负载——BERT、Llama2 LoRA、RetinaNet和Stable Diffusion——的训练时间数据,进行了详细分析,揭示了能够优化性能、GPU利用率与效率之间关系的配置方案。结果表明,在特定条件下存在一个临界点,可在降低训练时间的同时最大化整体效率。

原文摘要 · Abstract (English)

Training large-scale deep learning models has become a key challenge for the scientific community and industry. While the massive use of GPUs can significantly speed up training times, this approach has a negative impact on efficiency. In this article, we present a detailed analysis of the times reported by MLPerf Training v4.1 on four workloads: BERT, Llama2 LoRA, RetinaNet, and Stable Diffusion, showing that there are configurations that optimise the relationship between performance, GPU usage, and efficiency. The results point to a break-even point that allows training times to be reduced while maximising efficiency.

GPU优化训练效率AI硬件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。