arXiv:2411.10195cs.RO2024-11被引 8

用俯视图解决单目里程计尺度漂移问题,无需深度监督

BEV-ODOM: Reducing Scale Drift in Monocular Visual Odometry with BEV Representation

  • 引入俯视图表示,融合视角特征与相关性提取
  • 在多个数据集上显著降低长期运动估计的尺度漂移
  • 适合对精度要求高的自动驾驶与机器人导航场景

单目视觉里程计(MVO)在自主导航与机器人中至关重要,能提供低成本、灵活的运动追踪方案,但单目系统固有的尺度模糊性常导致累积误差。本文提出BEV-ODOM,一种基于鸟瞰图(BEV)表示的新型MVO框架,以缓解尺度漂移。不同于现有方法,BEV-ODOM集成深度引导的视角到俯视图编码器、相关性特征提取颈部及基于CNN-MLP的解码器,可在无深度监督或复杂优化的情况下实现三维运动估计(三个自由度)。该框架在长序列中有效减少尺度漂移,在NCLT、Oxford、KITTI等多个数据集上均实现高精度运动估计。实验表明,相较于现有MVO方法,BEV-ODOM在降低尺度漂移和提升精度方面表现更优。

原文摘要 · Abstract (English)

Monocular visual odometry (MVO) is vital in autonomous navigation and robotics, providing a cost-effective and flexible motion tracking solution, but the inherent scale ambiguity in monocular setups often leads to cumulative errors over time. In this paper, we present BEV-ODOM, a novel MVO framework leveraging the Bird's Eye View (BEV) Representation to address scale drift. Unlike existing approaches, BEV-ODOM integrates a depth-based perspective-view (PV) to BEV encoder, a correlation feature extraction neck, and a CNN-MLP-based decoder, enabling it to estimate motion across three degrees of freedom without the need for depth supervision or complex optimization techniques. Our framework reduces scale drift in long-term sequences and achieves accurate motion estimation across various datasets, including NCLT, Oxford, and KITTI. The results indicate that BEV-ODOM outperforms current MVO methods, demonstrating reduced scale drift and higher accuracy.

视觉里程计俯视图尺度漂移单目

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。