arXiv:2502.07600cs.CVcs.RO2025-02ICML被引 14

无需动作标注,从无标签视频中学习物体状态与隐式动作,实现可控视频预测。

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

  • 从无标注视频中联合推断物体表征与隐式动作
  • 在多个环境中超越基线模型的视频预测性能
  • 可结合用户输入或策略生成未来轨迹,适合机器人行为学习

预测未来场景表征是使机器人理解并交互环境的关键任务。然而,现有方法多依赖带精确动作标注的视频与仿真数据,难以利用大量无标注视频。为此,我们提出 PlaySlot——一种物体中心的视频预测模型,能从无标注视频序列中推断物体表征与隐式动作,并据此预测未来的物体状态与视频帧。该模型支持基于隐式动作生成多种可能的未来,这些动作可由视频动态推断、用户指定或由学习到的动作策略生成,从而实现灵活且可解释的世界建模。实验表明,PlaySlot 在不同环境下的视频预测性能优于随机与物体中心基线模型。此外,我们还证明所推断的隐式动作可用于高效地从无标注视频演示中学习机器人行为。视频与代码已公开于 https://play-slot.github.io/PlaySlot/。

原文摘要 · Abstract (English)

Predicting future scene representations is a crucial task for enabling robots to understand and interact with the environment. However, most existing methods rely on videos and simulations with precise action annotations, limiting their ability to leverage the large amount of available unlabeled video data. To address this challenge, we propose PlaySlot, an object-centric video prediction model that infers object representations and latent actions from unlabeled video sequences. It then uses these representations to forecast future object states and video frames. PlaySlot allows the generation of multiple possible futures conditioned on latent actions, which can be inferred from video dynamics, provided by a user, or generated by a learned action policy, thus enabling versatile and interpretable world modeling. Our results show that PlaySlot outperforms both stochastic and object-centric baselines for video prediction across different environments. Furthermore, we show that our inferred latent actions can be used to learn robot behaviors sample-efficiently from unlabeled video demonstrations. Videos and code are available on https://play-slot.github.io/PlaySlot/.

视频预测物体中心隐式动作机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。