arXiv:2411.12038cs.LGcs.AI2024-11被引 1

用Kubernetes在超算集群上自动扩展训练234个深度模型,提速科研效率。

Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster

  • 基于Kubernetes构建自动化训练流水线,实现分布式训练调度。
  • 在4,040小时内完成234个模型训练,涵盖遥感目标检测等三类应用。
  • 适合需要大规模模型实验的科研团队快速部署与管理训练任务。

在科学计算领域,深度学习算法在众多应用中表现出色。随着深度神经网络(DNN)不断发展,其训练所需的算力持续增长。如今,现代DNN需数百万次浮点运算,并耗时数天至数周才能完成训练。训练时间常成为各类深度学习研究的瓶颈,因此加速和扩展训练可显著提升研究效率。本文利用NRP Nautilus超算集群,为三个不同应用场景的DNN训练提供自动化与可扩展支持,包括遮挡物检测、烧毁区域分割和森林砍伐检测。总计在该集群上训练了234个深度神经网络模型,总训练时长达4,040小时。

原文摘要 · Abstract (English)

Throughout the scientific computing space, deep learning algorithms have shown excellent performance in a wide range of applications. As these deep neural networks (DNNs) continue to mature, the necessary compute required to train them has continued to grow. Today, modern DNNs require millions of FLOPs and days to weeks of training to generate a well-trained model. The training times required for DNNs are oftentimes a bottleneck in DNN research for a variety of deep learning applications, and as such, accelerating and scaling DNN training enables more robust and accelerated research. To that end, in this work, we explore utilizing the NRP Nautilus HyperCluster to automate and scale deep learning model training for three separate applications of DNNs, including overhead object detection, burned area segmentation, and deforestation detection. In total, 234 deep neural models are trained on Nautilus, for a total time of 4,040 hours

深度学习超算自动化训练遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。