用预训练深度模型提升360°立体匹配精度,显著改善复杂环境下的深度感知。
Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model
- 基于迭代优化框架,融合预训练单目相对深度特征
- 在真实数据集上将视差误差降低约16%
- 适合需要全向高精度深度感知的移动机器人场景
全景深度感知对需360°视野场景理解的移动机器人至关重要。基于相机的立体深度估计可低成本生成密集高分辨率深度图,无需依赖昂贵主动传感。然而,现有全景立体匹配方法在不同环境、深度范围和光照条件下深度精度有限,主要因真实世界数据稀缺。本文提出DFI-OmniStereo,一种新型全景立体匹配方法,利用大规模预训练基础模型提取相对单目深度特征,并嵌入基于迭代优化的立体匹配架构。通过专用两阶段训练策略,在尺度不变微调前充分利用相对单目深度特征。在真实世界Helvipad数据集上,该方法相较此前最佳全景立体匹配方法,视差平均绝对误差(MAE)降低约16%,达到当前最优水平。
原文摘要 · Abstract (English)
Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360° field of view. Camera-based setups offer a cost-effective option by using stereo depth estimation to generate dense, high-resolution depth maps without relying on expensive active sensing. However, existing omnidirectional stereo matching approaches achieve only limited depth accuracy across diverse environments, depth ranges, and lighting conditions, due to the scarcity of real-world data. We present DFI-OmniStereo, a novel omnidirectional stereo matching method that leverages a large-scale pre-trained foundation model for relative monocular depth estimation within an iterative optimization-based stereo matching architecture. We introduce a dedicated two-stage training strategy to utilize the relative monocular depth features for our omnidirectional stereo matching before scale-invariant fine-tuning. DFI-OmniStereo achieves state-of-the-art results on the real-world Helvipad dataset, reducing disparity MAE by approximately 16% compared to the previous best omnidirectional stereo method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。