让世界模型和动作模型在线协同进化,提升长程操作能力。
WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT

- 设计分层优化的强化学习框架,让世界模型与动作模型共同进化。
- 仅优化动作模型仅提升短程任务,联合优化才能显著改善长程任务表现。
- 首次将强化学习引入世界-动作模型,适合需要持续进化的机器人场景。
近期的世界-动作(WA)模型展现出强大的泛化能力和数据效率,但通常依赖专家轨迹进行训练。这种依赖限制了其在演示分布外获取精细操控技能的能力,并阻碍了通过真实环境交互持续改进。为解决这些问题,我们提出WAM-RL,一种通过与环境在线交互实现世界模型与动作模型联合优化的强化学习框架。通过让两个组件共同演化,该方法提升了精细控制与适应性。具体而言,一个WA模型包含世界模型和动作器(actor)。我们设计了定制化的强化学习方法,采用分层优化来协调两者的提升。方法层面,系统研究了对动作模型应用强化学习的影响,以及在强化学习设定下在线训练世界模型的效果。实验揭示关键洞察:仅优化动作器可提升短时任务性能,但在长时任务中无法带来显著收益;而联合优化世界模型与动作器是实现长程任务强表现的关键。本工作首次将强化学习引入世界-动作范式,为在线优化动作头与世界模型如何影响整体性能提供了重要见解。
原文摘要 · Abstract (English)
Recent World-Action (WA) models demonstrate strong generalization ability and data efficiency, but they typically rely on expert trajectories for training. This reliance limits their ability to acquire fine-grained manipulation skills beyond the demonstration distribution and prevents them from continuously improving through real-world interaction. To address these limitations, we propose WAM-RL, a reinforcement learning framework that enables joint optimization of the world model and the action model through online interaction with the environment. By allowing the two components to co-evolve, our approach enhances fine-grained control and adaptability. Specifically, a WA model consists of a world model and an actor. We design a tailored reinforcement learning method with hierarchical optimization to coordinate their improvement. On the methodological side, we systematically investigate the effects of applying reinforcement learning to the action model, as well as online training of the world model within an RL setting. Our experiments reveal a key insight: optimizing only the actor yields improvements on short-horizon tasks, but fails to provide significant gains on long-horizon tasks. In contrast, jointly optimizing both the world model and the actor is critical for achieving strong performance in long-horizon settings. Our work is the first to introduce reinforcement learning into the World-Action paradigm, and provides insights into how online optimization of both the action head and the world model impacts overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。