仅用物体3D模型,单图快速准确估计未知物体6自由度姿态。
Co-op: Correspondence-based Novel Object Pose Estimation
- 基于图像与预渲染模板的半密集对应关系,减少模板数量提升效率。
- 在BOP挑战7个核心数据集上显著超越现有方法,达到最新精度水平。
- 适合需要快速高精度姿态估计的工业检测与机器人抓取场景。
我们提出Co-op,一种从单张RGB图像中精确且鲁棒地估计训练时未见物体6自由度姿态的新方法。该方法仅需目标物体的CAD模型,无需额外微调即可精确估计姿态。相较于依赖大量模板导致效率低下的现有基于模型的方法,本方法通过在输入图像与预渲染模板间建立半密集对应关系,实现用少量模板即可快速准确估计。其性能提升得益于混合表示策略,结合了局部块分类与偏移回归。此外,姿态精修模块通过可微分PnP层,估计输入图像与渲染图像间的概率流,进一步优化初始估计。实验表明,该方法不仅推理速度快,且在BOP挑战赛的七个核心数据集上大幅领先于现有方法,达到当前最优性能。
原文摘要 · Abstract (English)
We propose Co-op, a novel method for accurately and robustly estimating the 6DoF pose of objects unseen during training from a single RGB image. Our method requires only the CAD model of the target object and can precisely estimate its pose without any additional fine-tuning. While existing model-based methods suffer from inefficiency due to using a large number of templates, our method enables fast and accurate estimation with a small number of templates. This improvement is achieved by finding semi-dense correspondences between the input image and the pre-rendered templates. Our method achieves strong generalization performance by leveraging a hybrid representation that combines patch-level classification and offset regression. Additionally, our pose refinement model estimates probabilistic flow between the input image and the rendered image, refining the initial estimate to an accurate pose using a differentiable PnP layer. We demonstrate that our method not only estimates object poses rapidly but also outperforms existing methods by a large margin on the seven core datasets of the BOP Challenge, achieving state-of-the-art accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。