无需对称标注,实现跨类别物体部件的高精度姿态估计
Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy

- 分阶段优化候选到最终四元数回归,提升姿态精度
- 将对称性预测建模为概率分布,自监督学习避免标注依赖
- 适用于数据稀缺场景,适合机器人交互与具身智能系统
高泛化能力的机器人物体交互与操作亟需高质量的跨类别物体感知。作为该领域先驱,可泛化且可操作的部件(GAParts)理解受到广泛关注。然而,现有方法或未充分考虑对称性问题,或依赖大量对称标注,严重限制了数据匮乏场景下的精准姿态估计。本文提出 SAFAG——一种无对称标注的通用可操作部件姿态估计框架。我们设计分阶段精炼的两阶段框架,实现从候选到最终四元数的回归;并将对称性预测建模为概率分布问题,采用自监督学习策略。实验表明,SAFAG 在性能与鲁棒性上均表现优异。我们认为该工作在具身智能系统中具有广泛应用潜力。
原文摘要 · Abstract (English)
Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actionable Parts (GAParts) understanding has attracted increasing attention from relevant researchers. However, most recent works either have insufficient design regarding the symmetry issue or require rich symmetry annotation, which severely impedes precise GAPart pose estimation in data-lacking scenarios. In this paper, we propose SAFAG, a novel Symmetry Annotation-Free framework for Generalizable and Actionable Parts Pose Estimation. Specifically, we suggest a stepwise refinement two-stage framework for candidate-to-final quaternion regression, and tackle the symmetry prediction as a probability distribution problem with self-supervised learning strategy. The experimental results demonstrate the superior performance and robustness of our SAFAG. We believe that our work has the enormous potential to be applied in many areas of embodied AI system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。