用双曲空间提升机器人抓取的3D视觉预训练效果
Hyperbolic Multiview Pretraining for Robotic Manipulation
- 采用双曲空间建模多视角3D特征,捕捉结构化关系
- 在多个任务和干扰条件下,性能优于主流基线方法
- 适合需要强空间感知的机器人抓取与操作研究
3D感知视觉预训练已被证明能有效提升下游机器人操作任务的性能。然而,现有方法受限于欧几里得嵌入空间,其平坦几何结构难以建模嵌入间的结构关系,导致难以学习对鲁棒空间感知至关重要的结构化表示。为此,我们提出HyperMVP——一种自监督的双曲多视角预训练框架。双曲空间具有捕捉复杂结构关系的几何优势。方法上,我们扩展了掩码自编码器范式,并设计了GeoLink编码器以学习多视角双曲表示。预训练编码器随后在操作任务上通过视觉-运动策略进行微调。此外,我们引入3D-MOV数据集,包含多种类型的3D点云,支持预训练。我们在COLOSSEUM、RLBench及真实场景中评估HyperMVP,结果表明其在多样任务与扰动设置下均持续优于强基线。实验验证了在非欧几里得空间中进行3D感知预训练,对学习鲁棒且可泛化的机器人操作策略具有潜力。
原文摘要 · Abstract (English)
3D-aware visual pretraining has proven effective in improving the performance of downstream robotic manipulation tasks. However, existing methods are constrained to Euclidean embedding spaces, whose flat geometry limits their ability to model structural relations among embeddings. As a result, they struggle to learn structured embeddings that are essential for robust spatial perception in robotic applications. To this end, we propose HyperMVP, a self-supervised framework for \underline{Hyper}bolic \underline{M}ulti\underline{V}iew \underline{P}retraining. Hyperbolic space offers geometric properties well suited for capturing structural relations. Methodologically, we extend the masked autoencoder paradigm and design a GeoLink encoder to learn multiview hyperbolic representations. The pretrained encoder is then finetuned with visuomotor policies on manipulation tasks. In addition, we introduce 3D-MOV, a large-scale dataset comprising multiple types of 3D point clouds to support pretraining. We evaluate HyperMVP on COLOSSEUM, RLBench, and real-world scenarios, where it consistently outperforms strong baselines across diverse tasks and perturbation settings. Our results highlight the potential of 3D-aware pretraining in a non-Euclidean space for learning robust and generalizable robotic manipulation policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。