用Transformer建模多物体多材料物理交互,直接从感知数据训练。
ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation
- 基于Transformer的点云世界模型,端到端学习多材质动态
- 在6个仿真+3个真实场景中预测误差更低,下游任务表现更优
- 无需复杂重建,适合真实机器人操控场景
3D世界模型通过学习环境演化中的物理规律,为可泛化的机器人操作提供新路径。但现有方法多局限于单材质动态,依赖耗时的3D场景重建来获取粒子轨迹。本文提出ParticleFormer,一种基于Transformer的点云世界模型,采用混合点云重建损失,同时监督全局与局部动态特征,在多材质、多物体机器人交互中实现细粒度建模。模型直接从真实机器人感知数据训练,无需复杂场景重建。我们在6个仿真和3个真实实验中验证其有效性,结果表明该模型在3D场景预测及下游模型预测控制(MPC)任务中均优于主流基线,动态预测精度更高,滚动误差更小。
原文摘要 · Abstract (English)
3D world models (i.e., learning-based 3D dynamics models) offer a promising approach to generalizable robotic manipulation by capturing the underlying physics of environment evolution conditioned on robot actions. However, existing 3D world models are primarily limited to single-material dynamics using a particle-based Graph Neural Network model, and often require time-consuming 3D scene reconstruction to obtain 3D particle tracks for training. In this work, we present ParticleFormer, a Transformer-based point cloud world model trained with a hybrid point cloud reconstruction loss, supervising both global and local dynamics features in multi-material, multi-object robot interactions. ParticleFormer captures fine-grained multi-object interactions between rigid, deformable, and flexible materials, trained directly from real-world robot perception data without an elaborate scene reconstruction. We demonstrate the model's effectiveness both in 3D scene forecasting tasks, and in downstream manipulation tasks using a Model Predictive Control (MPC) policy. In addition, we extend existing dynamics learning benchmarks to include diverse multi-material, multi-object interaction scenarios. We validate our method on six simulation and three real-world experiments, where it consistently outperforms leading baselines by achieving superior dynamics prediction accuracy and less rollout error in downstream visuomotor tasks. Experimental videos are available at https://suninghuang19.github.io/particleformer_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。