用神经运动模拟器实现高精度物理预测,让机器人零样本学新技能
Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
- 基于当前观测和动作预测系统未来物理状态
- 支持长时序精准预测,实现零样本强化学习
- 可将任意无模型RL算法转为模型驱动,提升采样效率
具身系统不仅需建模外部世界模式,还需理解自身运动动态。运动动态模型对高效技能习得与有效规划至关重要。本文提出神经运动模拟器(MoSim),一种基于当前观测与动作预测具身系统未来物理状态的世界模型。MoSim在物理状态预测上达到当前最佳性能,并在多种下游任务中表现优异。研究显示,当世界模型足够准确并能进行精确的长期预测时,可在虚拟世界中实现高效技能习得,甚至支持零样本强化学习。此外,MoSim可将任意无模型强化学习(RL)算法转化为模型基方法,有效解耦环境建模与RL算法开发。这一分离使得RL算法与世界建模可独立演进,显著提升样本效率并增强泛化能力。结果表明,针对运动动态的世界模型是发展更通用、更强具身系统的重要方向。
原文摘要 · Abstract (English)
An embodied system must not only model the patterns of the external world but also understand its own motion dynamics. A motion dynamic model is essential for efficient skill acquisition and effective planning. In this work, we introduce the neural motion simulator (MoSim), a world model that predicts the future physical state of an embodied system based on current observations and actions. MoSim achieves state-of-the-art performance in physical state prediction and provides competitive performance across a range of downstream tasks. This works shows that when a world model is accurate enough and performs precise long-horizon predictions, it can facilitate efficient skill acquisition in imagined worlds and even enable zero-shot reinforcement learning. Furthermore, MoSim can transform any model-free reinforcement learning (RL) algorithm into a model-based approach, effectively decoupling physical environment modeling from RL algorithm development. This separation allows for independent advancements in RL algorithms and world modeling, significantly improving sample efficiency and enhancing generalization capabilities. Our findings highlight that world models for motion dynamics is a promising direction for developing more versatile and capable embodied systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。