arXiv:2504.07955cs.CV2025-04ICCV被引 5

用物体角点提升稀疏视角下的姿态估计泛化能力

BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

  • 以物体边界框角点作为中间表示,建立2D-3D对应关系
  • 在遮挡和稀疏视图下仍能准确恢复角点,精度优于现有方法
  • 适合真实场景中未见过物体的姿态估计任务

本文提出一种基于RGB的通用物体姿态估计方法,针对稀疏视角场景下的挑战。现有方法虽能估计未见物体的姿态,但在遮挡和参考视图稀疏时泛化能力有限。为此,我们引入物体边界框的3D角点作为中间表示,其可从稀疏输入视图可靠恢复;同时设计了一种基于参考的点生成器,能在目标视图中估计出2D角点,即使在遮挡情况下也表现良好。这些角点作为语义点,可与3D角点建立2D-3D对应,用于PnP算法求解姿态。在YCB-Video和Occluded-LINEMOD数据集上的大量实验表明,本方法显著优于当前最优方法,验证了该表示的有效性及泛化能力的大幅提升,对实际应用至关重要。

原文摘要 · Abstract (English)

This paper presents a generalizable RGB-based approach for object pose estimation, specifically designed to address challenges in sparse-view settings. While existing methods can estimate the poses of unseen objects, their generalization ability remains limited in scenarios involving occlusions and sparse reference views, restricting their real-world applicability. To overcome these limitations, we introduce corner points of the object bounding box as an intermediate representation of the object pose. The 3D object corners can be reliably recovered from sparse input views, while the 2D corner points in the target view are estimated through a novel reference-based point synthesizer, which works well even in scenarios involving occlusions. As object semantic points, object corners naturally establish 2D-3D correspondences for object pose estimation with a PnP algorithm. Extensive experiments on the YCB-Video and Occluded-LINEMOD datasets show that our approach outperforms state-of-the-art methods, highlighting the effectiveness of the proposed representation and significantly enhancing the generalization capabilities of object pose estimation, which is crucial for real-world applications.

姿态估计稀疏视角角点表示泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。