单视角学习3D几何表征,提升机器人抓取的视点泛化能力。
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
- 仅用单视角图像预训练3D几何结构,避免多视角依赖。
- 在RLBench上平均成功率提升12.7%,视点变化下性能下降小于30%。
- 适合需要跨视角泛化的机器人抓取任务研究者使用。
现实世界中的机器人操作需要具备鲁棒的空间场景理解能力和强视点泛化能力的视觉运动策略。尽管近期3D感知视觉表征取得进展,但仍存在若干关键局限:推理时依赖多视角观测,不适用于单视角受限场景;场景建模不完整,难以捕捉精细几何结构以支持精准操作;缺乏有效策略训练方法来保留并利用获得的3D知识。为此,我们提出MethodName,一种统一的表征-策略学习框架,实现视点泛化的机器人操作。MethodName引入单视角3D预训练范式,通过点云重建与前向高斯喷溅,在多视角监督下学习整体几何表征。在策略学习阶段,采用多步蒸馏机制保留预训练几何理解,并有效迁移至操作技能。我们在12个RLBench任务上进行实验,方法平均成功率比当前最优方法提高12.7%。进一步在6个代表性任务上评估零样本视点泛化能力,中等和大角度视点偏移下的成功率下降分别为22.0%和29.7%,而当前最优方法分别下降41.6%和51.5%。
原文摘要 · Abstract (English)
Real-world robotic manipulation demands visuomotor policies capable of robust spatial scene understanding and strong generalization across diverse camera viewpoints. While recent advances in 3D-aware visual representations have shown promise, they still suffer from several key limitations, including reliance on multi-view observations during inference which is impractical in single-view restricted scenarios, incomplete scene modeling that fails to capture holistic and fine-grained geometric structures essential for precise manipulation, and lack of effective policy training strategies to retain and exploit the acquired 3D knowledge. To address these challenges, we present MethodName, a unified representation-policy learning framework for view-generalizable robotic manipulation. MethodName introduces a single-view 3D pretraining paradigm that leverages point cloud reconstruction and feed-forward gaussian splatting under multi-view supervision to learn holistic geometric representations. During policy learning, MethodName performs multi-step distillation to preserve the pretrained geometric understanding and effectively transfer it to manipulation skills. We conduct experiments on 12 RLBench tasks, where our approach outperforms the previous state-of-the-art method by 12.7% in average success rate. Further evaluation on six representative tasks demonstrates strong zero-shot view generalization, with success rate drops of only 22.0% and 29.7% under moderate and large viewpoint shifts respectively, whereas the state-of-the-art method suffers larger decreases of 41.6% and 51.5%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。