arXiv:2510.26170cs.ROcs.CV2025-10

融合全局与局部特征,提升单目相机在动态环境中的定位精度。

Self-localization on a 3D map by fusing global and local features from a monocular camera

  • 结合CNN局部特征与Vision Transformer全局特征进行定位
  • 动态障碍物下定位准确率是SOTA的1.5倍,误差降低20.1%
  • 实测机器人平均定位误差7.51cm,适合自动驾驶场景

基于廉价单目相机实现3D地图上的自定位是自动驾驶的关键需求。传统基于卷积神经网络(CNN)的方法依赖邻近像素提取局部特征,但在存在行人等动态障碍物时表现不佳。本文提出一种新方法,将擅长捕捉图像整体结构关系的Vision Transformer与CNN结合,同时利用全局与局部特征。实验表明,在含动态障碍物的CG数据集上,该方法的准确率提升幅度为无动态障碍物时的1.5倍;在公开数据集上,自定位误差比现有最先进方法(SOTA)减少20.1%。此外,使用该方法的机器人在实际测试中平均定位误差为7.51cm,优于当前SOTA。

原文摘要 · Abstract (English)

Self-localization on a 3D map by using an inexpensive monocular camera is required to realize autonomous driving. Self-localization based on a camera often uses a convolutional neural network (CNN) that can extract local features that are calculated by nearby pixels. However, when dynamic obstacles, such as people, are present, CNN does not work well. This study proposes a new method combining CNN with Vision Transformer, which excels at extracting global features that show the relationship of patches on whole image. Experimental results showed that, compared to the state-of-the-art method (SOTA), the accuracy improvement rate in a CG dataset with dynamic obstacles is 1.5 times higher than that without dynamic obstacles. Moreover, the self-localization error of our method is 20.1% smaller than that of SOTA on public datasets. Additionally, our robot using our method can localize itself with 7.51cm error on average, which is more accurate than SOTA.

自定位单目相机视觉里程计Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。