用数字孪生加速机器人视觉语言模型的现实世界学习效率。
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
- 构建数字孪生环境,通过虚拟训练扩展真实探索空间。
- 在虚拟环境中并行训练生成高质量轨迹,提升真实机器人学习速度。
- 仅需20分钟真实交互即实现近100%成功率,适合工业机器人部署。
尽管视觉-语言-动作(VLA)模型具备强泛化能力,但其仍受限于专家示范成本高和真实世界交互有限。在线强化学习虽有潜力,但在真实VLA操作中面临探索效率低、覆盖范围小的问题。通过系统性实验证明,线上RL的有效探索空间主要受监督微调(SFT)阶段轨迹分布的制约。为此,我们提出TwinRL——一种数字孪生与真实世界协同的后训练框架,包含三个阶段:SFT预热、孪生强化学习预热和真实世界强化学习。TwinRL首先从手机拍摄场景重建高保真数字孪生体。在SFT阶段,引入探索空间扩展策略,将轨迹分布支持范围扩展至真实示范之外,重塑探索空间以提高强化学习效率。不同于将孪生体视为数据增强工具,我们提出孪生强化学习预热策略,使其作为真实世界强化学习的探索引导器。具体而言,TwinRL在数字孪生中高效并行执行强化学习,生成交互轨迹填充经验回放池,并稳定后续真实世界学习。该过程还识别出易失败但信息丰富的配置,支持针对性的人机协同试错,进一步提升机器人效率。在四个任务上,TwinRL在分布内和分布外区域均实现接近100%的成功率,收敛速度比先前方法快30%以上,且仅需20分钟真实机器人交互。
原文摘要 · Abstract (English)
Despite strong generalization capabilities, Vision-Language-Action (VLA) models remain constrained by the high cost of expert demonstrations and limited real-world interaction. While online reinforcement learning (RL) has shown promise, its application to real-world VLA manipulation is hindered by low exploration efficiency and restricted exploration coverage. Through systematic real-world experiments, we observe that the effective exploration space of online RL is largely constrained by the trajectory distribution induced during supervised fine-tuning (SFT). Motivated by this observation, we propose TwinRL, a digital twin-real-world collaborative post-training framework that expands and guides RL exploration for VLA models through three stages: SFT warm-up, twin RL warm-up, and real-world RL. TwinRL first reconstructs a high-fidelity digital twin from smartphone-captured scenes. During the SFT stage, we introduce an exploration space expansion strategy that expands the support of the trajectory distribution beyond real demonstrations, reshaping the exploration space for more effective RL. Rather than treating the twin as a data augmentation tool, we propose a twin RL warm-up strategy that enables it to act as an exploration guide for real-world RL. Specifically, TwinRL performs efficient parallel RL in the digital twin to generate interactive trajectories that populate the replay buffer and stabilize subsequent real-world RL learning. This process also identifies failure-prone yet informative configurations, enabling targeted human-in-the-loop rollouts to further improve on-robot efficiency. Across four tasks, TwinRL achieves near-100% success in both in-distribution and out-of-distribution regions, delivering over 30% faster convergence than prior real-world RL methods with only 20 minutes of on-robot interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。