用合成数据训练视觉模型,实现火星探测车在空中地图中的精准定位。
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
- 采用双编码器网络实现地面与空中视角的跨视图定位。
- 在真实轨迹数据上定位误差小于1.5米,复杂路径也表现稳定。
- 适合对行星探测、无人机协同定位感兴趣的工程师和研究者。
行星机器人中的精确定位是支持未来任务大规模扩展的关键。英格拉姆直升机与多颗轨道器的成功为地面-空中机器人团队的应用奠定了基础。本文研究探测车如何仅通过有限视场的单目地面图像,在局部空域地图中实现自我定位。由于真实空间数据中带真值位置标签的数据稀缺,我们提出一种基于跨视图定位的双编码器深度神经网络方法。利用视觉基础模型进行语义分割,并结合大规模合成数据弥合真实图像与模拟数据之间的领域差距。我们还构建了一个新的真实世界探测车轨迹数据集(在行星类比设施中采集),包含对应的真实定位数据,以及大量对应的合成图像对。通过粒子滤波进行状态估计,仅依赖地面视角图像序列即可在简单与复杂轨迹上实现高精度定位。
原文摘要 · Abstract (English)
Accurate localisation in planetary robotics enables the advanced autonomy required to support the increased scale and scope of future missions. The successes of the Ingenuity helicopter and multiple planetary orbiters lay the groundwork for future missions that use ground-aerial robotic teams. In this paper, we consider rovers using machine learning to localise themselves in a local aerial map using limited field-of-view monocular ground-view RGB images as input. A key consideration for machine learning methods is that real space data with ground-truth position labels suitable for training is scarce. In this work, we propose a novel method of localising rovers in an aerial map using cross-view-localising dual-encoder deep neural networks. We leverage semantic segmentation with vision foundation models and high volume synthetic data to bridge the domain gap to real images. We also contribute a new cross-view dataset of real-world rover trajectories with corresponding ground-truth localisation data captured in a planetary analogue facility, plus a high volume dataset of analogous synthetic image pairs. Using particle filters for state estimation with the cross-view networks allows accurate position estimation over simple and complex trajectories based on sequences of ground-view images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。