千卡分布式训练平台让机器人智能训练提速40倍,突破大规模数据与算力瓶颈。
Thousand-GPU Large-Scale Training and Optimization Recipe for AI-Native Cloud Embodied Intelligence Infrastructure
- 基于LeRobot框架构建千卡云平台,重构数据流与训练流程
- 单轮训练从15小时缩短至22分钟,效率提升40倍;模型推理加速188%
- 适配大模型研发团队、具身智能系统开发者,推动人机协同落地
具身智能是迈向通用人工智能(AGI)的关键一步,但其发展面临数据、框架、基础设施和评估体系等多重挑战。本文首次在工业界推出基于云的千卡分布式训练平台,依托广泛采用的LeRobot框架,系统性攻克全流程瓶颈。在数据层,重构数据流水线以优化具身训练数据流转;训练层面,针对GR00T-N1.5模型,利用千万级数据与千卡集群,单轮训练时间由15小时降至22分钟,实现40倍加速;通过引入可变长FlashAttention与数据打包(Data Packing),实现从样本冗余到序列融合的转变,训练速度提升188%;π-0.5注意力优化使训练加速165%;FP8量化带来140%性能提升。基础设施方面,依托3.2T RDMA网络、高性能存储与基于Ray的弹性AI数据湖,实现数据、存储、通信与计算的深度协同。同时构建端到端评估系统,形成从训练、仿真到评估的闭环。该框架已在千卡集群上完成全面验证,为下一代自主智能机器人的研发与应用奠定关键技术基础,有望加速人机深度融合时代的到来。
原文摘要 · Abstract (English)
Embodied intelligence is a key step towards Artificial General Intelligence (AGI), yet its development faces multiple challenges including data, frameworks, infrastructure, and evaluation systems. To address these issues, we have, for the first time in the industry, launched a cloud-based, thousand-GPU distributed training platform for embodied intelligence, built upon the widely adopted LeRobot framework, and have systematically overcome bottlenecks across the entire pipeline. At the data layer, we have restructured the data pipeline to optimize the flow of embodied training data. In terms of training, for the GR00T-N1.5 model, utilizing thousand-GPU clusters and data at the scale of hundreds of millions, the single-round training time has been reduced from 15 hours to just 22 minutes, achieving a 40-fold speedup. At the model layer, by combining variable-length FlashAttention and Data Packing, we have moved from sample redundancy to sequence integration, resulting in a 188% speed increase; π-0.5 attention optimization has accelerated training by 165%; and FP8 quantization has delivered a 140% speedup. On the infrastructure side, relying on high-performance storage, a 3.2T RDMA network, and a Ray-driven elastic AI data lake, we have achieved deep synergy among data, storage, communication, and computation. We have also built an end-to-end evaluation system, creating a closed loop from training to simulation to assessment. This framework has already been fully validated on thousand-GPU clusters, laying a crucial technical foundation for the development and application of next-generation autonomous intelligent robots, and is expected to accelerate the arrival of the era of human-machine integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。