arXiv:2509.19713cs.CVcs.RO2025-09

用单目视觉惯性数据迭代优化像素尺度,实现高精度稀疏点深度估计。

VIMD: Monocular Visual-Inertial Motion and Depth Estimation

  • 通过多视角信息迭代修正每像素尺度,取代全局仿射建模。
  • 仅需10-20个稀疏度量点/图像即可实现高精度深度估计。
  • 模块化设计兼容多种深度模型,适合资源受限设备部署。

准确高效的稠密度量深度估计对机器人与扩展现实(XR)中的三维视觉感知至关重要。本文提出一种单目视觉惯性运动与深度(VIMD)学习框架,通过基于MSCKF的单目视觉惯性运动追踪,实现稠密度量深度估计。核心思想是利用多视角信息迭代优化每像素尺度,而非如以往工作采用全局不变仿射模型拟合。VIMD框架高度模块化,可适配多种现有深度估计主干网络。我们在TartanAir和VOID数据集上进行了广泛评估,并展示了在AR Table数据集上的零样本泛化能力。结果表明,即使仅有每图像10-20个稀疏度量深度点,VIMD仍能实现卓越的准确性和鲁棒性,使其成为资源受限场景下的实用解决方案,同时其强鲁棒性与泛化能力在多种应用场景中具有显著潜力。

原文摘要 · Abstract (English)

Accurate and efficient dense metric depth estimation is crucial for 3D visual perception in robotics and XR. In this paper, we develop a monocular visual-inertial motion and depth (VIMD) learning framework to estimate dense metric depth by leveraging accurate and efficient MSCKF-based monocular visual-inertial motion tracking. At the core the proposed VIMD is to exploit multi-view information to iteratively refine per-pixel scale, instead of globally fitting an invariant affine model as in the prior work. The VIMD framework is highly modular, making it compatible with a variety of existing depth estimation backbones. We conduct extensive evaluations on the TartanAir and VOID datasets and demonstrate its zero-shot generalization capabilities on the AR Table dataset. Our results show that VIMD achieves exceptional accuracy and robustness, even with extremely sparse points as few as 10-20 metric depth points per image. This makes the proposed VIMD a practical solution for deployment in resource constrained settings, while its robust performance and strong generalization capabilities offer significant potential across a wide range of scenarios.

深度估计视觉惯性单目稀疏点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。