解决机器人从堆叠物中选物并精准定位抓取的难题
JENGA: Object selection and pose estimation for robotic grasping from a stack
- 结合摄像头与惯性传感器,优先选择高层无遮挡物体
- 提出新数据集与评估指标,兼顾选物与姿态精度
- 在建筑场景砖块抓取中验证了实际可行性
基于视觉的机器人抓取通常针对孤立物体或散乱堆放的物体。然而,在建筑或仓储自动化等场景中,机器人需与堆叠结构交互。本文定义了从堆叠中选择适抓物体并估计其6自由度(6DoF)姿态的问题。为此,提出一种基于相机-惯性测量单元(camera-IMU)的方法,优先选取堆叠高层无遮挡物体,并构建一个用于基准测试的数据集及融合物体选择与姿态精度的评估指标。实验表明,尽管该方法表现良好,但实现完全无误的解决方案仍具挑战性。最后,在建筑场景的砖块抓取应用中展示了部署效果。
原文摘要 · Abstract (English)
Vision-based robotic object grasping is typically investigated in the context of isolated objects or unstructured object sets in bin picking scenarios. However, there are several settings, such as construction or warehouse automation, where a robot needs to interact with a structured object formation such as a stack. In this context, we define the problem of selecting suitable objects for grasping along with estimating an accurate 6DoF pose of these objects. To address this problem, we propose a camera-IMU based approach that prioritizes unobstructed objects on the higher layers of stacks and introduce a dataset for benchmarking and evaluation, along with a suitable evaluation metric that combines object selection with pose accuracy. Experimental results show that although our method can perform quite well, this is a challenging problem if a completely error-free solution is needed. Finally, we show results from the deployment of our method for a brick-picking application in a construction scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。