用单目深度模型提升立体匹配,零样本泛化更强。
DEFOM-Stereo: Depth Foundation Model Based Stereo Matching
- 融合单目相对深度模型与循环立体匹配框架。
- 在多个基准上达到最佳性能,部分指标排名第一。
- 适合需要强泛化能力的自动驾驶与机器人视觉场景。
立体匹配是计算机视觉与机器人中进行度量深度估计的关键技术。现实世界中的遮挡和无纹理区域会阻碍基于双目匹配线索的精确视差估计。近期,单目相对深度估计利用视觉基础模型展现出显著的泛化能力。为此,我们将在循环立体匹配框架中引入鲁棒的单目相对深度模型,构建基于深度基础模型的新型立体匹配框架——DEFOM-Stereo。在特征提取阶段,通过整合传统CNN与DEFOM的特征,构建联合上下文与匹配特征编码器;在更新阶段,使用DEFOM预测的深度初始化递归视差,并引入尺度更新模块以在正确尺度上优化视差。DEFOM-Stereo在零样本泛化能力上远超当前最优方法。此外,在KITTI 2012、KITTI 2015、Middlebury和ETH3D等多个基准测试中表现卓越,多项指标排名第一。在鲁棒视觉挑战的联合评估中,本模型同时优于以往各独立基准上的表现,进一步证明其强大性能。
原文摘要 · Abstract (English)
Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization using vision foundation models. Thus, to facilitate robust stereo matching with monocular depth cues, we incorporate a robust monocular relative depth model into the recurrent stereo-matching framework, building a new framework for depth foundation model-based stereo-matching, DEFOM-Stereo. In the feature extraction stage, we construct the combined context and matching feature encoder by integrating features from conventional CNNs and DEFOM. In the update stage, we use the depth predicted by DEFOM to initialize the recurrent disparity and introduce a scale update module to refine the disparity at the correct scale. DEFOM-Stereo is verified to have much stronger zero-shot generalization compared with SOTA methods. Moreover, DEFOM-Stereo achieves top performance on the KITTI 2012, KITTI 2015, Middlebury, and ETH3D benchmarks, ranking $1^{st}$ on many metrics. In the joint evaluation under the robust vision challenge, our model simultaneously outperforms previous models on the individual benchmarks, further demonstrating its outstanding capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。