arXiv:2603.20607cs.ROcs.LG2026-03被引 4

用虚拟世界模型训练机器人,减少真实交互成本。

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models

  • 用统一多模态模型高效构建虚拟环境。
  • 通过多视角交替解码保持视觉一致性。
  • 分块分支滚动降低误差累积,适合实际部署。

视觉-语言-动作(VLA)模型在机器人控制中展现出强大泛化能力,但通过强化学习(RL)微调受限于真实交互的高成本与安全风险。在可交互世界模型中训练VLA模型可避免这些问题,但也面临像素级建模、多视角一致性及稀疏奖励下误差累积等挑战。基于大模型与模型基于强化学习的最新进展,我们提出VLA-MBPO框架,针对这些难题设计三项关键策略:(i) 利用统一多模态模型(UMMs)实现数据高效的环境建模;(ii) 采用交错视角解码机制增强多视角一致性;(iii) 采用分块级分支滚动策略缓解误差传播。理论分析与仿真及真实任务实验表明,VLA-MBPO显著提升策略性能与样本效率,验证了其在真实机器人部署中的鲁棒性与可扩展性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety risks of real-world interaction. Training VLA models in interactive world models avoids these issues but introduces several challenges, including pixel-level world modeling, multi-view consistency, and compounding errors under sparse rewards. Building on recent advances across large multimodal models and model-based RL, we propose VLA-MBPO, a practical framework to tackle these problems in VLA finetuning. Our approach has three key design choices: (i) adapting unified multimodal models (UMMs) for data-efficient world modeling; (ii) an interleaved view decoding mechanism to enforce multi-view consistency; and (iii) chunk-level branched rollout to mitigate error compounding. Theoretical analysis and experiments across simulation and real-world tasks demonstrate that VLA-MBPO significantly improves policy performance and sample efficiency, underscoring its robustness and scalability for real-world robotic deployment.

机器人控制世界模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。