arXiv:2503.03726cs.CVcs.RO2025-03被引 7

通过多视角图像和主动视觉,精准估计无纹理物体的6自由度位姿。

Active 6D Pose Estimation for Textureless Objects using Multi-View RGB Frames

  • 分两步:先估平移,再用标准模板匹配定方向。
  • 在ROBI、TOD等数据集上精度超越现有方法,且所需视角更少。
  • 适合需要高精度位姿估计的机器人抓取与操作任务。

从RGB图像中估计无纹理物体的6D位姿是机器人领域的重要问题。由于外观模糊、旋转对称性和严重遮挡,单视角方法难以应对广泛物体,促使研究向多视角位姿估计和下一最佳视角预测发展。本文提出一种基于纯RGB图像的主动感知框架,核心思想是将6D位姿估计分解为两阶段顺序过程:首先估计3D平移,解决RGB图像固有的尺度与深度模糊;随后利用标准尺度模板匹配简化3D方向估计。在此基础上,引入主动感知策略,预测下一最佳相机视角,有效降低位姿不确定性并提升精度。在公开数据集ROBI、TOD及自建透明物体数据集T-ROBI上评估,相同视角下性能显著优于现有方法;结合下一最佳视角策略,在所有数据集上以更少视角实现更高精度。相关视频与数据集将发布于项目主页:https://trailab.github.io/ActiveODPE。

原文摘要 · Abstract (English)

Estimating the 6D pose of textureless objects from RGB images is an important problem in robotics. Due to appearance ambiguities, rotational symmetries, and severe occlusions, single-view based 6D pose estimators are still unable to handle a wide range of objects, motivating research towards multi-view pose estimation and next-best-view prediction that addresses these limitations. In this work, we propose a comprehensive active perception framework for estimating the 6D poses of textureless objects using only RGB images. Our approach is built upon a key idea: decoupling the 6D pose estimation into a two-step sequential process can greatly improve both accuracy and efficiency. First, we estimate the 3D translation of each object, resolving scale and depth ambiguities inherent to RGB images. These estimates are then used to simplify the subsequent task of determining the 3D orientation, which we achieve through canonical scale template matching. Building on this formulation, we then introduce an active perception strategy that predicts the next best camera viewpoint to capture an RGB image, effectively reducing object pose uncertainty and enhancing pose accuracy. We evaluate our method on the public ROBI and TOD datasets, as well as on our reconstructed transparent object dataset, T-ROBI. Under the same camera viewpoints, our multi-view pose estimation significantly outperforms state-of-the-art approaches. Furthermore, by leveraging our next-best-view strategy, our approach achieves high pose accuracy with fewer viewpoints than heuristic-based policies across all evaluated datasets. The accompanying video and T-ROBI dataset will be released on our project page: https://trailab.github.io/ActiveODPE.

6D位姿主动感知多视角机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。