融合几何与语义,提升机器人操作的3D感知能力
CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment
- 结合掩码自编码与跨模态对比学习,兼顾全局结构与局部细节
- 在仿真和真实场景中显著优于现有方法,提升视觉运动策略性能
- 适合需要精细3D操作的机器人研究者与工程师
3D点云中的空间信息对机器人操作至关重要。然而,现有的3D预训练方法存在根本性权衡:掩码自编码(MAE)擅长捕捉空间几何特征但缺乏语义,而对比学习虽能从2D基础模型中提取语义,却难以处理操作任务所需的细粒度细节。为此,我们提出CLAR,一种新型3D预训练框架,通过统一MAE与全局跨模态对比学习,实现稳健的空间感知与丰富语义理解的融合。为增强对细粒度细节的关注,在局部层面引入可变形注意力驱动的自适应对齐机制,强制建立局部3D几何与2D视觉特征间的精确对应,克服传统全局对齐在操作任务中的局限。大量仿真与真实世界实验表明,CLAR在视觉运动策略学习中达到当前最优性能。
原文摘要 · Abstract (English)
The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fundamental trade-off: Masked Autoencoding (MAE) excels at capturing spatial-geometric features but lacks semantics, whereas contrastive learning, while able to distill semantics from 2D foundation models, is ill-suited for the fine-grained details required for manipulation tasks. To address these challenges, we propose CLAR, a novel 3D pre-training framework that synergizes global understanding with fine-grained local alignment. Our framework unifies MAE with global cross-modal contrastive learning to integrate robust spatial awareness with rich semantic understanding. To enhance its focus on fine-grained details, at the local level, we introduce an adaptive alignment mechanism that leverages deformable attention to force precise correspondences between local 3D geometry and 2D visual features, thereby overcoming the limitations of conventional global alignment in manipulation tasks. Extensive experiments in simulation and the real world demonstrate that CLAR achieves state-of-the-art performance, significantly outperforming existing methods in visuomotor policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。