用3D多视角对比学习提升机器人抓取的精度与泛化能力
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
- 通过点云重建多视角图像,融合深度与3D坐标信息
- 在大规模模拟轨迹上用对比学习关联3D结构与动作模式
- 适配真实场景,显著提升小样本任务的训练效率与性能
在行为克隆策略中利用预训练的2D图像表示已取得显著成功,但这类表示难以捕捉物体与场景的3D空间信息,影响精确操作。本文提出一种新型3D预训练框架CLAMP,基于点云与机器人动作,从RGB-D图像和相机外参合并生成点云,并重渲染包含动态腕部视角在内的多视角四通道图像观测,以更清晰呈现目标物体。通过在大规模模拟机器人轨迹上进行对比学习,预训练编码器学会将物体的3D几何与位置信息与机器人动作模式对齐。预训练阶段还使用扩散策略初始化策略权重,以提升微调时的样本效率和性能。微调阶段仅需少量任务示范即可获得良好表现。实验表明,该设计大幅提升了未见任务的学习效率与策略性能,在六个模拟任务和五个真实任务中均优于现有先进方法。
原文摘要 · Abstract (English)
Leveraging pre-trained 2D image representations in behavior cloning policies has achieved great success and has become a standard approach for robotic manipulation. However, such representations fail to capture the 3D spatial information about objects and scenes that is essential for precise manipulation. In this work, we introduce Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining (CLAMP), a novel 3D pre-training framework that utilizes point clouds and robot actions. From the merged point cloud computed from RGB-D images and camera extrinsics, we re-render multi-view four-channel image observations with depth and 3D coordinates, including dynamic wrist views, to provide clearer views of target objects for high-precision manipulation tasks. The pre-trained encoders learn to associate the 3D geometric and positional information of objects with robot action patterns via contrastive learning on large-scale simulated robot trajectories. During encoder pre-training, we pre-train a Diffusion Policy to initialize the policy weights for fine-tuning, which is essential for improving fine-tuning sample efficiency and performance. After pre-training, we fine-tune the policy on a limited amount of task demonstrations using the learned image and action representations. We demonstrate that this pre-training and fine-tuning design substantially improves learning efficiency and policy performance on unseen tasks. Furthermore, we show that CLAMP outperforms state-of-the-art baselines across six simulated tasks and five real-world tasks. The project website and videos can be found at https://clamp3d.github.io/CLAMP/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。