用单目相机图像投影到球面,实现3D旋转对称的机器人操作策略学习。
3D Equivariant Visuomotor Policy Learning via Spherical Projection
- 将2D图像特征投影到球面,隐式建模SO(3)对称性,无需点云重建。
- 在仿真和真实场景中均显著提升性能与样本效率,超越强基线。
- 首个仅用单目RGB输入的SO(3)等变机器人策略框架,适合视觉感知任务。
等变模型已被证明能显著提升扩散策略的数据效率。然而,已有工作主要聚焦于多相机固定布置生成的点云输入,这不适用于当前主流的机载RGB相机(如GoPro)设置。本文通过在扩散策略模型中引入将2D RGB图像特征投影到球面的过程,填补了这一空白。该方法使我们能在不显式重建点云的情况下,对SO(3)对称性进行推理。我们在仿真和真实世界中进行了大量实验,结果表明该方法在性能和样本效率上均持续优于强基线。我们的工作——图像到球面策略(Image-to-Sphere Policy, ISP),是首个仅使用单目RGB输入的SO(3)等变机器人操作策略学习框架。
原文摘要 · Abstract (English)
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused primarily on point cloud inputs generated by multiple cameras fixed in the workspace. This type of point cloud input is not compatible with the now-common setting where the primary input modality is an eye-in-hand RGB camera like a GoPro. This paper closes this gap by incorporating into the diffusion policy model a process that projects features from the 2D RGB camera image onto a sphere. This enables us to reason about symmetries in $\mathrm{SO}(3)$ without explicitly reconstructing a point cloud. We perform extensive experiments in both simulation and the real world that demonstrate that our method consistently outperforms strong baselines in terms of both performance and sample efficiency. Our work, Image-to-Sphere Policy ($\textbf{ISP}$), is the first $\mathrm{SO}(3)$-equivariant policy learning framework for robotic manipulation that works using only monocular RGB inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。