arXiv:2602.02454cs.ROcs.AI2026-02被引 22

用世界模型训练机器人,比传统方法快18倍,还能在新场景中自我优化。

World-Gymnast: Training Robots with Reinforcement Learning in a World Model

  • 在视频世界模型中滚动执行策略,用视觉语言模型评分反馈。
  • 在Bridge机器人上性能比监督微调高18倍,比软件仿真高2倍。
  • 支持语言指令泛化、新场景测试时训练,适合云端机器人研发者。

真实世界中的机器人学习受限于物理交互成本。现有方法如从专家示范中监督微调(SFT)和在软件模拟器中强化学习(RL),分别受专家数据量和仿真到现实的差距限制。随着从真实视频-动作数据中学习的世界模型出现,我们探索在世界模型中训练策略是否比传统方法更有效。提出World-Gymnast:通过动作条件视频世界模型进行策略滚动,并用视觉语言模型(VLM)评估奖励,实现视觉-语言-动作(VLA)策略的强化学习微调。在Bridge机器人设置下,World-Gymnast性能比SFT提升最高达18倍,比软件模拟器提升最高达2倍。更重要的是,该方法展现出多项能力:支持多样语言指令、在世界模型生成的新场景中训练、在新场景中进行测试时训练,以及在线迭代优化世界模型与策略。结果表明,在云端构建世界模型并训练机器人策略,可能是弥合演示型机器人与家庭通用机器人之间差距的关键。

原文摘要 · Abstract (English)

Robot learning from interacting with the physical world is fundamentally bottlenecked by the cost of physical interaction. The two alternatives, supervised finetuning (SFT) from expert demonstrations and reinforcement learning (RL) in a software-based simulator, are limited by the amount of expert data available and the sim-to-real gap for manipulation. With the recent emergence of world models learned from real-world video-action data, we ask the question of whether training a policy in a world model can be more effective than supervised learning or software simulation in achieving better real-robot performance. We propose World-Gymnast, which performs RL finetuning of a vision-language-action (VLA) policy by rolling out the policy in an action-conditioned video world model and rewarding the rollouts with a vision-language model (VLM). On the Bridge robot setup, World-Gymnast outperforms SFT by as much as 18x and outperforms software simulator by as much as 2x. More importantly, World-Gymnast demonstrates intriguing capabilities of RL with a world model, including training on diverse language instructions and novel scenes from the world model, test-time training in a novel scene, and online iterative world model and policy improvement. Our results suggest learning a world model and training robot policies in the cloud could be the key to bridging the gap between robots that work in demonstrations and robots that can work in anyone's household.

强化学习世界模型机器人VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。