无需标注,自动从视频中识别物体并建模其动态行为。
Latent Particle World Models: Self-supervised Object-centric Stochastic Dynamics Modeling
- 从视频中自监督学习物体关键点与掩码,无需人工标注。
- 在真实与合成数据集上达到当前最佳的随机动态建模效果。
- 可灵活用于动作、语言或图像目标条件下的决策任务。
我们提出潜粒子世界模型(LPWM),一种自监督、以物体为中心的世界模型,可扩展至真实世界多物体数据集并应用于决策任务。LPWM能直接从视频数据中自主发现关键点、边界框和物体掩码,实现无需监督的丰富场景分解。其架构完全端到端训练,仅依赖视频输入,并支持对动作、语言和图像目标的灵活条件化。通过创新的潜行动模块,LPWM对随机粒子动态进行建模,在多种真实世界与合成数据集上取得当前最优表现。除了随机视频建模外,本文还展示了其在目标条件化模仿学习中的直接应用。代码、数据、预训练模型及视频回放已公开:https://taldatech.github.io/lpwm-web
原文摘要 · Abstract (English)
We introduce Latent Particle World Model (LPWM), a self-supervised object-centric world model scaled to real-world multi-object datasets and applicable in decision-making. LPWM autonomously discovers keypoints, bounding boxes, and object masks directly from video data, enabling it to learn rich scene decompositions without supervision. Our architecture is trained end-to-end purely from videos and supports flexible conditioning on actions, language, and image goals. LPWM models stochastic particle dynamics via a novel latent action module and achieves state-of-the-art results on diverse real-world and synthetic datasets. Beyond stochastic video modeling, LPWM is readily applicable to decision-making, including goal-conditioned imitation learning, as we demonstrate in the paper. Code, data, pre-trained models and video rollouts are available: https://taldatech.github.io/lpwm-web
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。