arXiv:2505.18474cs.RO2025-05被引 12

用统一的3D表示提升机器人视觉模仿学习的泛化能力

Canonical Policy: Learning Canonical 3D Representation for SE(3)-Equivariant Policy

  • 构建3D规范表示理论,实现不同视角下点云的统一映射
  • 在仿真和真实场景中分别提升18.0%和39.7%的性能表现
  • 适合需要强泛化能力的机器人操控任务研究者参考

视觉模仿学习在机器人操作中取得显著进展,但对未见物体、场景布局和相机视角的泛化仍是关键挑战。现有方法通过使用3D点云提供几何感知、外观不变的表示,并引入等变性以利用空间对称性。然而,现有等变方法常因组件非结构化整合而缺乏可解释性和严谨性。本文提出规范策略(Canonical Policy),一种面向3D等变模仿学习的原理性框架,通过统一3D点云观测到规范表示。我们首先建立3D规范表示理论,使可观测点云(包括新出现的)能被归入同一规范形式,从而实现等变的观测-动作映射。随后设计灵活的策略学习流程,融合规范表示中的几何对称性与现代生成模型的表达能力。我们在12个模拟任务和4个真实世界操作任务共16种配置下验证该方法,涵盖物体颜色、形状、相机视角及机器人平台的变化。相比最先进的模仿学习策略,该方法在仿真中平均提升18.0%,在真实实验中提升39.7%,展现出更优的泛化能力和样本效率。

原文摘要 · Abstract (English)

Visual Imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy architectures to exploit spatial symmetries. However, existing equivariant approaches often lack interpretability and rigor due to unstructured integration of equivariant components. We introduce canonical policy, a principled framework for 3D equivariant imitation learning that unifies 3D point cloud observations under a canonical representation. We first establish a theory of 3D canonical representations, enabling equivariant observation-to-action mappings by grouping both seen and novel point clouds to a canonical representation. We then propose a flexible policy learning pipeline that leverages geometric symmetries from canonical representation and the expressiveness of modern generative models. We validate canonical policy on 12 diverse simulated tasks and 4 real-world manipulation tasks across 16 configurations, involving variations in object color, shape, camera viewpoint, and robot platform. Compared to state-of-the-art imitation learning policies, canonical policy achieves an average improvement of 18.0% in simulation and 39.7% in real-world experiments, demonstrating superior generalization capability and sample efficiency. For more details, please refer to the project website: https://zhangzhiyuanzhang.github.io/cp-website/.

机器人学习3D泛化等变网络模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。