让机器人边部署边学习,持续提升通用操作能力。
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

- 通过真实机器人集群的部署数据,实现从离线到在线的持续强化学习。
- 16台双臂机器人在8个任务上平均成功率达95%,长时任务提升最显著。
- 适合需要长期优化、多任务泛化的机器人系统研发人员。
通用机器人策略越来越依赖大规模预训练,但仅靠离线数据难以应对真实世界中的分布偏移、长尾故障、任务变化和人工干预机会。本文提出“边部署边学习”(LWD)框架,实现基于视觉-语言-动作(VLA)策略的舰队级离线到在线强化学习。从预训练的VLA策略出发,利用整个机器人车队自主执行与人工干预收集的数据,形成部署、共享经验、策略优化、再部署的闭环。为稳定处理异构、稀疏奖励的舰队数据,LWD结合分布隐式价值学习(DIVL)进行鲁棒价值估计,以及通过伴随匹配的Q学习(QAM)提取流模型生成的动作策略。在16台双臂机器人组成的车队上,对8个真实场景下的操作任务进行了验证,包括语义商品补货及持续3–5分钟的长时任务。单一通用策略随车队经验积累不断改进,平均成功率达95%,长时任务收益最为明显。
原文摘要 · Abstract (English)
Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. Deployed robots encounter distribution shifts, long-tail failures, task variations, and human correction opportunities that fixed demonstration datasets cannot fully capture. We present Learning While Deploying (LWD), a fleet-scale offline-to-online reinforcement learning framework for continual post-training of generalist Vision-Language-Action (VLA) policies. Starting from a pretrained VLA policy, LWD closes the loop between deployment, shared physical experience, policy improvement, and redeployment by using autonomous rollouts and human interventions collected across a robot fleet. To stabilize learning from heterogeneous, sparse-reward fleet data, LWD combines Distributional Implicit Value Learning (DIVL) for robust value estimation with Q-learning via Adjoint Matching (QAM) for policy extraction in flow-based VLA action generators. We validate LWD on a fleet of 16 dual-arm robots across eight real-world manipulation tasks, including semantic grocery restocking and 3--5 minute long-horizon tasks. A single generalist policy improves as fleet experience accumulates, reaching an average success rate of 95%, with the largest gains on long-horizon tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。