仅用一张图像即可精准估算未知物体的6D位姿与尺寸,无需3D模型。
Any6D: Model-free 6D Pose Estimation of Novel Objects
- 通过联合对齐提升2D-3D匹配与尺度估计精度。
- 在五大数据集上超越当前最优方法,尤其在遮挡和光照变化下表现优异。
- 适合需要快速部署、无预训练3D模型的工业场景应用。
我们提出Any6D,一种无需3D模型的6D物体位姿估计框架,仅需一张RGB-D锚图即可在新场景中估计未知物体的6D位姿与尺寸。不同于依赖纹理3D模型或多视角输入的方法,Any6D采用联合对齐机制,增强2D-3D匹配与度量尺度估计,从而提升位姿精度。该方法结合渲染对比策略生成并优化位姿假设,在遮挡、非重叠视角、多变光照及跨环境差异等复杂条件下仍具鲁棒性。我们在五个挑战性数据集(REAL275、Toyota-Light、HO3D、YCBINEOAT、LM-O)上评估,结果表明其显著优于现有最先进方法。项目页面:https://taeyeop.com/any6d
原文摘要 · Abstract (English)
We introduce Any6D, a model-free framework for 6D object pose estimation that requires only a single RGB-D anchor image to estimate both the 6D pose and size of unknown objects in novel scenes. Unlike existing methods that rely on textured 3D models or multiple viewpoints, Any6D leverages a joint object alignment process to enhance 2D-3D alignment and metric scale estimation for improved pose accuracy. Our approach integrates a render-and-compare strategy to generate and refine pose hypotheses, enabling robust performance in scenarios with occlusions, non-overlapping views, diverse lighting conditions, and large cross-environment variations. We evaluate our method on five challenging datasets: REAL275, Toyota-Light, HO3D, YCBINEOAT, and LM-O, demonstrating its effectiveness in significantly outperforming state-of-the-art methods for novel object pose estimation. Project page: https://taeyeop.com/any6d
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。