arXiv:2409.10441cs.ROcs.CV2024-09ICRA被引 23

解决机器人部分可见时的位姿估计难题,提升实际操作鲁棒性。

CtRNet-X: Camera-to-Robot Pose Estimation in Real-world Conditions Using a Single Camera

  • 用视觉-语言模型精准定位机器人零部件,支持部分可见场景
  • 在公开数据集和自建部分视角数据上均实现稳定精度
  • 适合真实场景下机器人视觉控制,尤其适用于遮挡频繁的任务

相机到机器人的标定对基于视觉的机器人控制至关重要,传统方法需耗时的物理设置。近年来无标记位姿估计方法无需复杂布置即可达到高精度,但依赖所有机械臂关节始终在视野内。现实中,机器人常进出画面,部分结构可能全程不可见,导致特征不足而失效。为此,本文提出新框架,在部分可见条件下仍可估计机器人位姿。方法利用视觉-语言模型进行细粒度部件检测,并集成至关键点位姿网络,增强在多种工况下的鲁棒性。在多个公开数据集及自建的部分视图数据集上验证了其泛化能力与稳定性。结果表明,该方法能有效应对更广泛的真实操作场景中的位姿估计挑战。

原文摘要 · Abstract (English)

Camera-to-robot calibration is crucial for vision-based robot control and requires effort to make it accurate. Recent advancements in markerless pose estimation methods have eliminated the need for time-consuming physical setups for camera-to-robot calibration. While the existing markerless pose estimation methods have demonstrated impressive accuracy without the need for cumbersome setups, they rely on the assumption that all the robot joints are visible within the camera's field of view. However, in practice, robots usually move in and out of view, and some portion of the robot may stay out-of-frame during the whole manipulation task due to real-world constraints, leading to a lack of sufficient visual features and subsequent failure of these approaches. To address this challenge and enhance the applicability to vision-based robot control, we propose a novel framework capable of estimating the robot pose with partially visible robot manipulators. Our approach leverages the Vision-Language Models for fine-grained robot components detection, and integrates it into a keypoint-based pose estimation network, which enables more robust performance in varied operational conditions. The framework is evaluated on both public robot datasets and self-collected partial-view datasets to demonstrate our robustness and generalizability. As a result, this method is effective for robot pose estimation in a wider range of real-world manipulation scenarios.

位姿估计视觉-语言机器人控制部分可见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。