用高斯点云建模刚体运动,实现动作条件下的三维环境预测。
Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects
- 以对象为中心的高斯表示法,支持任意形状与多物体场景建模。
- 在合成数据上实现高精度未来运动预测,误差低于1.2cm。
- 适合需要精准物理推理的机器人非抓取操作任务。
世界模型使智能体能够预测自身动作对环境的影响。本文提出多刚体高斯世界模型(MRO-GWM),学习三维环境中刚体对象的动作条件动态。通过对象中心的高斯表示,可建模任意物体形状和多物体场景。我们设计了一种新颖的时空变压器架构,从历史对象高斯和未来动作中预测刚体运动。物体在标准坐标系中由其高斯表示,使运动描述为刚体变换。模型在多视角重建数据上训练,需处理因遮挡导致的部分观测。我们在包含典型家用物品的合成数据集上评估了该方法在多物体动态与交互中的预测性能,实验使用机器人末端执行器进行测试。此外,还在仿真环境中评估了该模型在非抓取操作中的模型预测控制表现。
原文摘要 · Abstract (English)
World models enable intelligent agents to predict the consequences of their actions on the environment. In this paper, we propose Multi Rigid Object Gaussian World Model (MRO-GWM), a novel model that learns action-conditional dynamics of rigid objects in 3D. By representing the scene by object-centric Gaussians, we can represent arbitrary object shapes and multi-object scenes. We develop a novel spatio-temporal transformer architecture that predicts future rigid body motion from a history of object Gaussians and future actions. Objects are represented by their Gaussians in a canonical frame, which allows for describing object motion as rigid body transformation. Our model is trained on reconstructions from multiple viewpoints, which requires the model to handle partial observations of objects due to occlusions. We analyze prediction performance of our approach on synthetic datasets composed of typical household objects with multi-object dynamics and interactions by a robot end effector. We also evaluate our model in model-predictive control for non-prehensile manipulation in simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。