提出旋转等变的体积抓取模型,提升采样效率与实时性能。
Equivariant Volumetric Grasping
- 采用三平面特征表示,水平面特征对90°旋转等变
- 结合流匹配生成抓取方向,实测在真实场景中性能领先
- 兼顾计算与内存开销,适合机器人抓取部署
我们提出一种新型体积抓取模型,对垂直轴旋转具有等变性,显著提升采样效率。模型采用三平面体积特征表示——即3D特征在三个标准平面上的投影。我们设计了一种新式三平面特征:水平面上的特征对90°旋转保持等变,而另外两平面特征之和对相同变换保持不变。我们进一步为两种先进体积抓取规划器(GIGA 和 IGD)开发了等变适配版本,具体包括推导出变形注意力机制的新等变形式,并提出基于流匹配的等变抓取方向生成模型。我们提供了所提等变性质的详细分析证明,并通过大量仿真与真实世界实验验证方法有效性。结果表明,该投影式设计降低了计算与内存开销;基于三平面特征构建的等变抓取模型始终优于非等变基线,在实时成本约束下实现更高性能。
原文摘要 · Abstract (English)
We propose a new volumetric grasp model that is equivariant to rotations around the vertical axis, leading to a significant improvement in sampling efficiency. Our model employs a tri-plane volumetric feature representation -- i.e., the projection of 3D features onto three canonical planes. We introduce a novel tri-plane feature design in which features on the horizontal plane are \emph{equivariant} to $90^\circ$ rotations, while the \emph{sum} of features from the other two planes remains \emph{invariant} to reflections induced by the same transformations. We further develop equivariant adaptations of two state-of-the-art volumetric grasp planners, GIGA and IGD. Specifically, we derive a new equivariant formulation of IGD's deformable attention mechanism and propose an equivariant generative model of grasp orientations based on flow matching. We provide a detailed analytical justification of the proposed equivariance properties and validate our approach through extensive simulated and real-world experiments. Our results demonstrate that the proposed projection-based design reduces both computational and memory costs. Moreover, the equivariant grasp models built on top of our tri-plane features consistently outperform their non-equivariant counterparts, achieving higher performance within a real-time cost constraint. Video and code can be viewed in: https://mousecpn.github.io/evg-page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。