arXiv:2503.13926cs.CV2025-03ICLR被引 11

用球面表示学习形状无关的变换,提升类别级物体姿态估计精度

Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose Estimation

  • 以球面作为统一代理形状,实现跨形态物体的通用坐标映射
  • 在多个基准上达到新最好结果,显著提升对应点预测精度
  • 适合做类别级三维视觉任务的研究者和工业应用开发者

类别级物体姿态估计旨在确定特定类别中新物体的姿态与尺寸。现有基于对应关系的方法通常采用点表示来建立观测点与归一化物体坐标间的对应。然而,由于标准坐标的固有形状依赖性,这些方法在不同物体形状间存在语义不一致问题。为此,我们创新性地利用球面作为物体的共享代理形状,通过球面表示学习形状无关的变换。基于此,提出新架构SpherePose,通过三项核心设计实现精准对应点预测:第一,点特征提取具有SO(3)不变性,可在任意旋转下稳健映射相机空间与物体空间;第二,球面注意力机制从全局视角传播和融合球面锚点特征,缓解噪声与点云不完整的影响;第三,设计双曲对应损失函数,有效区分细微差异,提升对应预测精度。在CAMERA25、REAL275和HouseCat6D三个基准上的实验表明,该方法性能优越,验证了球面表示与架构创新的有效性。

原文摘要 · Abstract (English)

Category-level object pose estimation aims to determine the pose and size of novel objects in specific categories. Existing correspondence-based approaches typically adopt point-based representations to establish the correspondences between primitive observed points and normalized object coordinates. However, due to the inherent shape-dependence of canonical coordinates, these methods suffer from semantic incoherence across diverse object shapes. To resolve this issue, we innovatively leverage the sphere as a shared proxy shape of objects to learn shape-independent transformation via spherical representations. Based on this insight, we introduce a novel architecture called SpherePose, which yields precise correspondence prediction through three core designs. Firstly, We endow the point-wise feature extraction with SO(3)-invariance, which facilitates robust mapping between camera coordinate space and object coordinate space regardless of rotation transformation. Secondly, the spherical attention mechanism is designed to propagate and integrate features among spherical anchors from a comprehensive perspective, thus mitigating the interference of noise and incomplete point cloud. Lastly, a hyperbolic correspondence loss function is designed to distinguish subtle distinctions, which can promote the precision of correspondence prediction. Experimental results on CAMERA25, REAL275 and HouseCat6D benchmarks demonstrate the superior performance of our method, verifying the effectiveness of spherical representations and architectural innovations.

姿态估计球面表示对应点预测三维视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。