无需姿态标签,通过识别机器人上的LED状态实现单目位姿估计。
Self-supervised Learning Of Visual Pose Estimation Without Pose Labels By Classifying LED States
- 用LED状态分类任务自监督训练,无需姿态标签或机器人模型。
- 在真实场景中随机移动两台机器人采集数据,不依赖外部设备。
- 可推广至不同场景,支持多机器人同时定位,性能媲美有监督方法。
我们提出一种单目RGB相对位姿估计算法,从零开始训练,无需姿态标签或对机器人形状、外观的先验知识。训练时假设:(i) 机器人配备多个状态独立且已知的LED;(ii) 每个LED的大致观测方向已知;(iii) 有一张已知目标距离的标定图像,以解决单目深度估计的模糊性。训练数据由两台机器人随机运动生成,无需外部基础设施或人工干预。模型通过预测图像中每个LED的状态进行训练,从而学习到机器人的图像位置、距离和相对方位。推理时,LED状态未知且任意,不影响位姿估计性能。定量实验表明,该方法在性能上可与依赖姿态标签或CAD模型的先进方法相媲美,具备良好的跨域泛化能力,并支持多机器人位姿估计。
原文摘要 · Abstract (English)
We introduce a model for monocular RGB relative pose estimation of a ground robot that trains from scratch without pose labels nor prior knowledge about the robot's shape or appearance. At training time, we assume: (i) a robot fitted with multiple LEDs, whose states are independent and known at each frame; (ii) knowledge of the approximate viewing direction of each LED; and (iii) availability of a calibration image with a known target distance, to address the ambiguity of monocular depth estimation. Training data is collected by a pair of robots moving randomly without needing external infrastructure or human supervision. Our model trains on the task of predicting from an image the state of each LED on the robot. In doing so, it learns to predict the position of the robot in the image, its distance, and its relative bearing. At inference time, the state of the LEDs is unknown, can be arbitrary, and does not affect the pose estimation performance. Quantitative experiments indicate that our approach: is competitive with SoA approaches that require supervision from pose labels or a CAD model of the robot; generalizes to different domains; and handles multi-robot pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。