arXiv:2602.06508cs.RO2026-02被引 27

用闭环训练提升视觉语言动作模型在真实世界中的表现

World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

  • 构建包含成功与近成功轨迹的SANS数据集,提升动作与结果对齐
  • 联合预测未来帧和二值奖励,使奖励估计更可靠
  • 通过虚拟环境迭代优化模型和策略,减少真实交互成本

强化学习(RL)可超越行为克隆改进视觉-语言-动作(VLA)策略,但真实世界中的强化学习因大量试错、重置、监督和安全风险而代价高昂。行动条件视频世界模型可在虚拟环境中训练,但其动作跟随能力不精确,尤其在细微近成功失败场景中表现差,且缺乏原生奖励信号。基于不准确视觉预测计算奖励不可靠。本文提出World-VLA-Loop,包含两项基础设计和一个协同演化范式:首先构建SANS数据集,专门混合成功与近成功轨迹以改善动作-结果对齐;其次训练一个状态感知的视频世界模型,从扩散潜变量联合预测未来帧与二值奖励,将奖励估计耦合至生成器而非独立模块,从而反哺视觉预测。由于VLA行为在强化学习中动态变化,固定模拟器可能与更新后的策略失配,因此World-VLA-Loop通过闭环机制,使用优化后的世界模型进行迭代式VLA后训练,并将每次策略的回放数据反馈用于增强和微调世界模型。在仿真与真实机器人实验中,该方法显著提升VLA性能,同时大幅降低对昂贵物理交互的依赖。

原文摘要 · Abstract (English)

Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to extensive rollouts, resets, supervision, and safety risks. Action-conditioned video world models offer an option to train in virtual environments, yet they exhibit imprecise action following, particularly on subtle near-success failures. Besides, they lack native reward signals for RL. Computing rewards based on inaccurate visual predictions remain unreliable. We introduce World-VLA-Loop, structured around two foundational designs and a higher-level co-evolving paradigm. We first curate SANS, dedicatedly mixing successful and near-success trajectories to improve action-outcome alignment. Then, we train a state-aware video world model that jointly predicts future frames and binary rewards from diffusion latents. It couples reward estimation to the generator rather than a separate module, and in turn, benefits visual prediction. Since VLA behavior shifts during RL, a fixed simulator can misalign with the updated policy, World-VLA-Loop therefore closes the loop by using the refined world model for iterative VLA post-training while feeding rollouts from each improved policy back to augment and fine-tune the world model. Across simulation and real-robot experiments, World-VLA-Loop substantially improves VLA performance while reducing reliance on costly physical interaction.

强化学习视频建模闭环训练真实世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。