让机器人通过实时协作学习,快速提升任务能力。
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
- 构建闭环系统,多机器人实时上传经验,云端同步更新策略。
- 数小时内完成有效训练,机器人越多,性能提升越接近线性增长。
- 支持多种算法,适合需要快速适应真实场景的通用机器人系统。
视觉-语言-动作(VLA)模型通过大规模预训练实现强泛化能力,但实际部署需达到专家级任务表现。现有后训练方法多为离线、单机或任务特定,难以实现高效在线适应与规模化学习。本文提出可扩展的在线后训练系统(SOP),支持在物理世界中对通用VLA模型进行分布式、多任务的在线后训练。SOP通过闭环架构将机器人集群持续流式传输在线经验与人工干预信号至中心云学习器,并异步接收更新策略。该设计支持即时在线修正,通过并行部署扩大数据收集规模,并在适应过程中保持模型泛化性。SOP与后训练算法解耦,我们采用交互式模仿学习(HG-DAgger)和强化学习(RECAP)两种实例。在布料折叠、盒子组装、杂货补货等真实操作任务中,SOP显著提升了大模型性能,且维持单一共享策略。仅需数小时真实交互即可实现有效后训练,性能随机器人数量近似线性提升。结果表明,将在线学习与大规模部署紧密结合,是实现高效、可靠、可扩展的通用机器人策略后训练的关键。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models achieve strong generalization through large-scale pre-training, but real-world deployment requires expert-level task proficiency in addition to broad generality. Existing post-training approaches for VLA models are typically offline, single-robot, or task-specific, limiting effective on-policy adaptation and scalable learning from real-world interaction. We introduce a Scalable Online Post-training (SOP) system that enables online, distributed, multi-task post-training of generalist VLA models directly in the physical world. SOP tightly couples execution and learning through a closed-loop architecture in which a fleet of robots continuously streams on-policy experience and human intervention signals to a centralized cloud learner, and asynchronously receives updated policies. This design supports prompt on-policy correction, scales experience collection through parallel deployment, and preserves generality during adaptation. SOP is agnostic to the choice of post-training algorithm; we instantiate it with both interactive imitation learning (HG-DAgger) and reinforcement learning (RECAP). Across a range of real-world manipulation tasks including cloth folding, box assembly, and grocery restocking, we show that SOP substantially improves the performance of large pretrained VLA models while maintaining a single shared policy across tasks. Effective post-training can be achieved within hours of real-world interaction, and performance scales near-linearly with the number of robots in the fleet. These results suggest that tightly coupling online learning with fleet-scale deployment is instrumental to enabling efficient, reliable, and scalable post-training of generalist robot policies in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。