arXiv:2511.09515cs.ROcs.AI2025-11被引 50

用视觉语言动作模型自建世界模型,实现高效自学习机器人操控

WMPO: World Model-based Policy Optimization for Vision-Language-Action Models

  • 基于像素预测构建世界模型,让虚拟试错更贴近真实感知
  • 在仿真中仅用少量交互数据就达到超越传统方法的性能
  • 适合需要自我纠错和长期学习的复杂机器人任务

视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但其依赖专家示范,难以从失败中学习并自我修正。强化学习(RL)可通过与物理环境交互实现自我提升,但在真实机器人上样本效率极低。我们提出基于世界模型的策略优化(WMPO),一种无需与真实环境交互的在线策略VLA强化学习框架。不同于广泛使用的潜在世界模型,WMPO聚焦于像素级预测,使“想象”轨迹与利用网络规模图像预训练的VLA特征对齐。关键在于,WMPO支持在线策略的GRPO,性能显著优于常用的离线策略方法。大量仿真与真实机器人实验表明,WMPO(i)大幅提高样本效率,(ii)实现更强整体表现,(iii)涌现出自我纠正等新行为,(iv)具备鲁棒泛化与持续学习能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, but their reliance on expert demonstrations limits their ability to learn from failures and perform self-corrections. Reinforcement learning (RL) addresses these through self-improving interactions with the physical environment, but suffers from high sample complexity on real robots. We introduce World-Model-based Policy Optimization (WMPO), a principled framework for on-policy VLA RL without interacting with the real environment. In contrast to widely used latent world models, WMPO focuses on pixel-based predictions that align the "imagined" trajectories with the VLA features pretrained with web-scale images. Crucially, WMPO enables the policy to perform on-policy GRPO that provides stronger performance than the often-used off-policy methods. Extensive experiments in both simulation and real-robot settings demonstrate that WMPO (i) substantially improves sample efficiency, (ii) achieves stronger overall performance, (iii) exhibits emergent behaviors such as self-correction, and (iv) demonstrates robust generalization and lifelong learning capabilities.

机器人操控强化学习世界模型自学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。