arXiv:2412.11530cs.CV2024-12被引 4

用预训练深度模型提升单目视觉里程计的鲁棒性和度量精度

RoMeO: Robust Metric Visual Odometry

  • 融合单目深度与多视图立体模型,实现度量尺度估计
  • 在6个数据集上相对和绝对轨迹误差降低超50%
  • 适合野外复杂场景的机器人与AR/VR应用

视觉里程计(VO)旨在从视觉输入中估计相机位姿,是VR/AR和机器人等应用的核心。本文聚焦单目RGB VO,输入仅为无IMU或3D传感器的单目视频。现有方法在此挑战性场景下缺乏鲁棒性,难以泛化到未见数据(尤其是户外),且无法恢复度量尺度。我们提出鲁棒度量视觉里程计(RoMeO),利用预训练深度模型的先验信息解决上述问题。RoMeO结合单目度量深度与多视图立体(MVS)模型,实现度量尺度估计,简化特征匹配,改善初始化并正则化优化。通过训练时注入噪声及自适应过滤噪声深度先验,确保在真实世界数据上的鲁棒性。如图1所示,RoMeO在覆盖室内外场景的6个多样化数据集上显著超越现有最优方法(DPVO),相对和绝对轨迹误差均降低超过50%。性能提升亦可迁移至完整SLAM系统(含全局束调整与回环检测)。代码将在论文接受后公开。

原文摘要 · Abstract (English)

Visual odometry (VO) aims to estimate camera poses from visual inputs -- a fundamental building block for many applications such as VR/AR and robotics. This work focuses on monocular RGB VO where the input is a monocular RGB video without IMU or 3D sensors. Existing approaches lack robustness under this challenging scenario and fail to generalize to unseen data (especially outdoors); they also cannot recover metric-scale poses. We propose Robust Metric Visual Odometry (RoMeO), a novel method that resolves these issues leveraging priors from pre-trained depth models. RoMeO incorporates both monocular metric depth and multi-view stereo (MVS) models to recover metric-scale, simplify correspondence search, provide better initialization and regularize optimization. Effective strategies are proposed to inject noise during training and adaptively filter noisy depth priors, which ensure the robustness of RoMeO on in-the-wild data. As shown in Fig.1, RoMeO advances the state-of-the-art (SOTA) by a large margin across 6 diverse datasets covering both indoor and outdoor scenes. Compared to the current SOTA DPVO, RoMeO reduces the relative (align the trajectory scale with GT) and absolute trajectory errors both by >50%. The performance gain also transfers to the full SLAM pipeline (with global BA & loop closure). Code will be released upon acceptance.

视觉里程计度量恢复单目感知鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。