首个基于点云的3D机器人操作基础模型,仅需80次演示即能高效泛化。
FP3: A 3D Foundation Policy for Robotic Manipulation
- 采用扩散Transformer架构,用6万条点云轨迹预训练
- 80次示范后在新环境新物体上达90%以上成功率
- 首次实现3D几何信息驱动的机器人策略泛化,适合具身智能研究者
继自然语言处理与计算机视觉之后,大规模多任务数据预训练的基础模型在机器人领域也展现出巨大潜力。然而,现有机器人基础模型多依赖2D图像观测,忽略了对机器人感知和理解三维世界至关重要的3D几何信息。本文提出FP3,首个用于机器人操作的大规模3D基础策略模型。FP3基于可扩展的扩散Transformer架构,在包含60,000条轨迹的点云观测数据上进行预训练。得益于模型设计与多样化预训练数据,FP3可在下游任务中高效微调并表现出强大的泛化能力。真实机器人实验表明,仅需80次示范,FP3即可在包含未见物体的新环境中完成新任务,成功率超过90%,显著超越现有机器人基础模型。
原文摘要 · Abstract (English)
Following its success in natural language processing and computer vision, foundation models that are pre-trained on large-scale multi-task datasets have also shown great potential in robotics. However, most existing robot foundation models rely solely on 2D image observations, ignoring 3D geometric information, which is essential for robots to perceive and reason about the 3D world. In this paper, we introduce FP3, a first large-scale 3D foundation policy model for robotic manipulation. FP3 builds on a scalable diffusion transformer architecture and is pre-trained on 60k trajectories with point cloud observations. With the model design and diverse pre-training data, FP3 can be efficiently fine-tuned for downstream tasks while exhibiting strong generalization capabilities. Experiments on real robots demonstrate that with only 80 demonstrations, FP3 is able to learn a new task with over 90% success rates in novel environments with unseen objects, significantly surpassing existing robot foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。