arXiv:2411.18623cs.CV2024-11被引 57

将2D大模型升级为3D机器人操作能力,提升空间感知与抓取精度。

Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation

  • 用掩码自编码器让2D模型学会隐式理解3D空间关系
  • 通过位置映射实现点云直接输入,减少几何信息损失
  • 在仿真和真实场景中均超越现有方法,适合复杂抓取任务

3D几何信息对机器人操作至关重要,需感知环境、推理空间关系并应对复杂配置。当前研究虽重视显式3D特征提取,但仍受限于缺乏大规模机器人3D数据及空间几何信息丢失问题。为此,我们提出Lift3D框架,逐步增强2D基础模型的隐式与显式3D机器人表征,构建鲁棒的3D操作策略。首先设计任务感知掩码自编码器,掩蔽与任务相关的可操作区域并重建深度信息,提升2D模型的隐式3D表征能力。经自监督微调后,引入2D模型升维策略,建立输入3D点与2D模型位置嵌入间的映射关系。基于此映射,Lift3D可直接使用2D基础模型编码点云数据,利用大规模预训练知识构建显式3D机器人表征,同时最小化空间信息损失。实验表明,Lift3D在多个仿真基准和真实场景中持续优于此前最先进方法。

原文摘要 · Abstract (English)

3D geometric information is essential for manipulation tasks, as robots need to perceive the 3D environment, reason about spatial relationships, and interact with intricate spatial configurations. Recent research has increasingly focused on the explicit extraction of 3D features, while still facing challenges such as the lack of large-scale robotic 3D data and the potential loss of spatial geometry. To address these limitations, we propose the Lift3D framework, which progressively enhances 2D foundation models with implicit and explicit 3D robotic representations to construct a robust 3D manipulation policy. Specifically, we first design a task-aware masked autoencoder that masks task-relevant affordance patches and reconstructs depth information, enhancing the 2D foundation model's implicit 3D robotic representation. After self-supervised fine-tuning, we introduce a 2D model-lifting strategy that establishes a positional mapping between the input 3D points and the positional embeddings of the 2D model. Based on the mapping, Lift3D utilizes the 2D foundation model to directly encode point cloud data, leveraging large-scale pretrained knowledge to construct explicit 3D robotic representations while minimizing spatial information loss. In experiments, Lift3D consistently outperforms previous state-of-the-art methods across several simulation benchmarks and real-world scenarios.

3D操作2D模型点云编码机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。