arXiv:2603.09030cs.ROcs.AI2026-03被引 11

用机器人自玩数据训练高保真世界模型,提升物理交互预测能力

PlayWorld: Learning Robot World Models from Autonomous Play

  • 通过机器人自主玩耍收集数据,无需人类示范
  • 在接触丰富的操作任务中实现更真实的物理交互预测
  • 适合需要真实物理模拟的机器人强化学习与故障预测场景

动作条件视频模型为构建可直接从数据中学习的通用机器人模拟器提供了有前景的路径。然而,尽管在大规模机器人数据集上训练,当前最先进的视频模型仍难以预测对机器人操作至关重要的物理一致的机器人-物体交互。为弥合这一差距,我们提出 PlayWorld,一种简单、可扩展且完全自治的管道,用于从交互经验中训练高保真视频世界模拟器。与依赖成功导向的人类示范的先前方法不同,PlayWorld 是首个完全从无监督机器人自玩数据中学习的系统,实现了自然可扩展的数据收集,同时捕捉了建模真实物体动力学所必需的复杂、长尾物理交互。在多种操纵任务上的实验表明,PlayWorld 能生成高质量、物理一致的接触丰富交互预测,这是人类收集数据训练的世界模型所无法捕捉的。我们进一步展示了 PlayWorld 在细粒度故障预测和策略评估中的多功能性,性能相比人类数据提升高达 40%。最后,我们证明了 PlayWorld 可支持在世界模型内进行强化学习,部署到真实世界后策略成功率提升 65%。

原文摘要 · Abstract (English)

Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still struggle to predict physically consistent robot-object interactions that are crucial in robotic manipulation. To close this gap, we present PlayWorld, a simple, scalable, and fully autonomous pipeline for training high-fidelity video world simulators from interaction experience. In contrast to prior approaches that rely on success-biased human demonstrations, PlayWorld is the first system capable of learning entirely from unsupervised robot self-play, enabling naturally scalable data collection while capturing complex, long-tailed physical interactions essential for modeling realistic object dynamics. Experiments across diverse manipulation tasks show that PlayWorld generates high-quality, physically consistent predictions for contact-rich interactions that are not captured by world models trained on human-collected data. We further demonstrate the versatility of PlayWorld in enabling fine-grained failure prediction and policy evaluation, with up to 40% improvements over human-collected data. Finally, we demonstrate how PlayWorld enables reinforcement learning in the world model, improving policy performance by 65% in success rates when deployed in the real world.

世界模型机器人自玩强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。