arXiv:2512.20538cs.CV2025-12

多视角特征对齐实现无需训练的6D姿态估计

AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

  • 通过多视角特征-度量优化,统一求解世界坐标系下的物体姿态
  • 在6个数据集上超越现有方法,工业场景下表现尤为突出
  • 无需特定物体训练或对称标注,通用性强适合实际部署

单视角RGB基于模型的物体6D姿态估计方法虽具强泛化能力,但受限于深度模糊、杂乱与遮挡。多视角方法有望解决这些问题,但现有工作依赖精确的单视角姿态估计或缺乏对未见物体的泛化能力。本文提出AlignPose,一种利用多个外参标定的RGB视角进行6D姿态估计的方法,无需任何物体特定训练或对称性标注。其核心是专为物体姿态设计的多视角特征-度量精炼机制,通过同时最小化实时渲染物体特征与各视角观测特征间的差异,优化单一一致的世界坐标系姿态。在六个数据集(YCB-V、T-LESS、HouseCat6D、ITODD-MV、IPD、XYZ-IBD)上使用BOP基准评估,结果表明AlignPose优于其他已发表方法,尤其在多视角易获取的工业数据集上表现显著。

原文摘要 · Abstract (English)

Single-view RGB model-based object pose estimation methods achieve strong generalization but are fundamentally limited by depth ambiguity, clutter, and occlusions. Multi-view pose estimation methods have the potential to solve these issues, but existing works rely on precise single-view pose estimates or lack generalization to unseen objects. We address these challenges via the following three contributions. First, we introduce AlignPose, a 6D object pose estimation method that aggregates information from multiple extrinsically calibrated RGB views and does not require any object-specific training or symmetry annotation. Second, the key component of this approach is a new multi-view feature-metric refinement specifically designed for object pose. It optimizes a single, consistent world-frame object pose by minimizing the feature discrepancy between on-the-fly rendered object features and observed image features across all views simultaneously. Third, we report extensive experiments on six datasets (YCB-V, T-LESS, HouseCat6D, ITODD-MV, IPD, XYZ-IBD) using the BOP benchmark evaluation and show that AlignPose outperforms other published methods, especially on challenging industrial datasets where multiple views are readily available in practice.

6D姿态估计多视角特征对齐工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。