用生成式3D世界实现机器人视觉语言动作模型的高效仿真到现实迁移。
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
- 基于生成式3D世界和语言驱动场景设计,自动生成海量多样交互场景。
- 仿真成功率从9.7%提升至79.8%,真实世界成功率从21.7%升至75%。
- 适合追求高泛化能力与低成本数据扩展的机器人强化学习研究者。
大型视觉语言模型(VLM)在强化学习(RL)中的优异表现激发了将类似方法用于机器人视觉语言动作(VLA)模型微调的研究。许多近期工作直接在真实世界中微调VLA以避免仿真到现实的差距。然而,真实世界强化学习固有地限制了VLA的泛化能力,因物理世界中大规模扩展场景与物体多样性成本过高。这导致预训练模型反而演变为过拟合、场景依赖的策略。通过仿真训练可获取多样化场景,但场景构建成本高昂。本文提出利用生成式3D世界模型,结合语言驱动场景设计,生成数百个包含独特物体与背景的交互场景,实现可扩展且高度并行的策略学习。从预训练模仿学习基线出发,本方法使仿真成功率从9.7%提升至79.8%,任务完成时间提速1.25倍。进一步验证了生成数字孪生质量与领域随机化的协同效应,真实世界成功率从21.7%提升至75%,任务速度提升1.13倍。最后,消融实验表明,增加场景多样性可直接提升零样本泛化能力,凸显生成式3D数据的无限潜力。
原文摘要 · Abstract (English)
The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs directly in the real world to avoid addressing the sim-to-real gap. While real-world RL circumvents sim-to-real issues, it inherently limits the generality of the resulting VLA, as scaling scene and object diversity in the physical world is prohibitively difficult. This leads to the paradoxical outcome of transforming a broadly pretrained model into an overfitted, scene-specific policy. Training in simulation can instead provide access to diverse scenes, but designing those scenes is also costly. In this work, we show that VLAs can be RL fine-tuned without sacrificing generality and with reduced labor by leveraging 3D world generative models. Using these models together with a language-driven scene designer, we generate hundreds of diverse interactive scenes containing unique objects and backgrounds, enabling scalable and highly parallel policy learning. Starting from a pretrained imitation baseline, our approach increases simulation success from 9.7% to 79.8% while achieving a 1.25$\times$ speedup in task completion time. We further demonstrate successful sim-to-real transfer enabled by the quality of the generated digital twins together with domain randomization, improving real-world success from 21.7% to 75% and achieving a 1.13$\times$ speedup. Finally, we further highlight the benefits of leveraging the effectively unlimited data from 3D world generative models through an ablation study showing that increasing scene diversity directly improves zero-shot generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。