arXiv:2602.11323cs.CV2026-02

用学习到的深度先验提升单目惯导在低纹理环境下的定位精度。

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

  • 将深度先验融入VINS-Mono优化框架,保持边缘设备实时性。
  • 在TartanGround和M3ED数据集上降低28.3%的轨迹误差。
  • 适合做边缘部署的高精度视觉惯性定位系统研究者参考。

传统单目视觉惯性里程计(VIO)在低纹理环境中因视觉特征稀疏而难以准确估计位姿。为解决此问题,密集单目深度估计(MDE)被广泛探索作为补充信息源。尽管基于视觉变压器(ViT)的复杂基础模型可提供稠密且几何一致的深度,但其计算开销通常无法满足实时边缘部署需求。本文通过将学习到的深度先验直接集成至VINS-Mono优化后端,提出一种新框架,强制执行仿射不变深度一致性与成对序数约束,并通过方差门控显式过滤不稳定伪影。该方法严格遵守边缘设备的计算限制,同时稳健恢复度量尺度。在TartanGround和M3ED数据集上的大量实验表明,该方法可在挑战性场景中防止发散,并显著提升精度,绝对轨迹误差(ATE)最高降低28.3%。代码将公开。

原文摘要 · Abstract (English)

Traditional monocular Visual-Inertial Odometry (VIO) systems struggle in low-texture environments where sparse visual features are insufficient for accurate pose estimation. To address this, dense Monocular Depth Estimation (MDE) has been widely explored as a complementary information source. While recent Vision Transformer (ViT) based complex foundational models offer dense, geometrically consistent depth, their computational demands typically preclude them from real-time edge deployment. Our work bridges this gap by integrating learned depth priors directly into the VINS-Mono optimization backend. We propose a novel framework that enforces affine-invariant depth consistency and pairwise ordinal constraints, explicitly filtering unstable artifacts via variance-based gating. This approach strictly adheres to the computational limits of edge devices while robustly recovering metric scale. Extensive experiments on the TartanGround and M3ED datasets demonstrate that our method prevents divergence in challenging scenarios and delivers significant accuracy gains, reducing Absolute Trajectory Error (ATE) by up to 28.3%. Code will be made available.

视觉惯性深度估计边缘计算位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。