用可控世界模型让机器人在虚拟环境里稳定练技能,避免幻觉误导。
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
- 通过可控动作条件视频模型提升虚拟推演稳定性
- 关键帧初始化推演减少误差累积,实现长程可靠模拟
- 模型与策略共同进化,适合想用强化学习优化机器人的研究者
强化学习(RL)有望突破视觉-语言-动作(VLA)模型的模仿学习局限,但其对大量真实世界交互的需求阻碍了物理机器人直接部署。近期工作尝试用学习到的世界模型作为模拟器进行策略优化,但闭环想象推演不可避免地出现幻觉和长时序误差累积,不仅降低视觉质量,还因不可靠的学习信号误导策略优化。我们提出WoVR,一种基于可靠世界模型的后训练VLA策略强化学习框架。不同于假设世界模型完全准确,WoVR显式调控强化学习与不完美想象动态的交互方式:通过可控制的动作条件视频世界模型提升推演稳定性;借助关键帧初始化推演降低有效误差深度;通过世界模型与策略的协同演化保持一致性。大量实验表明,WoVR实现了稳定的长周期想象推演与有效的策略优化,在LIBERO基准上表现更优,并在多个机器人平台上均获得一致的现实性能提升。结果表明,当幻觉被显式控制时,世界模型可作为实际可用的强化学习模拟器。更多可视化内容见https://wovr-corl.github.io。
原文摘要 · Abstract (English)
Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots. Recent work attempts to use learned world models as simulators for policy optimization, yet closed-loop imagined rollouts inevitably suffer from hallucination and long-horizon error accumulation. Such errors not only degrade visual fidelity, but also mislead policy optimization by providing unreliable learning signals. We propose WoVR, a reliable world-model-based RL framework for post-training VLA policies. Instead of assuming a faithful world model, WoVR explicitly regulates how RL interacts with imperfect imagined dynamics. It improves rollout stability through a controllable action-conditioned video world model, reshapes imagined interaction to reduce effective error depth via Keyframe-Initialized Rollouts, and maintains policy--simulator alignment through World Model-Policy co-evolution. Extensive experiments demonstrate that WoVR enables stable long-horizon imagined rollouts and effective policy optimization, achieving superior LIBERO performance and consistent real-world gains across multiple robotic platforms. These results show that world models can serve as practical simulators for RL when hallucination is explicitly controlled. Additional visualization results are available at https://wovr-corl.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。