无需参考视图的全景立体匹配,通过多视角一致性提升深度估计精度。
Reference-Free Omnidirectional Stereo Matching via Multi-View Consistency Maximization
- 基于视角间相关性建模,不依赖特定参考视图
- 在多个数据集上实现更一致、更鲁棒的全景深度估计
- 适合机器人导航等需要全局感知的应用场景
从多鱼眼立体匹配中可靠地进行全景深度估计对许多应用(如具身机器人)至关重要。现有方法或依赖球面扫描与启发式融合构建代价列,或基于校正视图的参考中心立体匹配,但均未能显式利用多视角间的几何关系,难以捕捉全局依赖、可见性或尺度变化。本文提出一种新范式——无参考框架FreeOmniMVS,通过多视角一致性最大化实现深度估计。其核心是将成对相关性聚合为鲁棒、可见性感知且全局一致的共识,对遮挡、部分重叠和基线变化具有强容错能力。具体地,引入视图对相关性变压器(VCT),显式建模所有相机视图对之间的相关性体积,可剔除因遮挡或离焦导致不可靠的视图对;同时设计轻量级注意力机制,自适应融合相关向量,无需指定参考视图,使所有摄像头平等参与匹配过程。大量实验表明,该方法在多种基准数据集上实现了全局一致、可见性感知和尺度感知的全景深度估计性能领先。
原文摘要 · Abstract (English)
Reliable omnidirectional depth estimation from multi-fisheye stereo matching is pivotal to many applications, such as embodied robotics. Existing approaches either rely on spherical sweeping with heuristic fusion strategies to build the cost columns or perform reference-centric stereo matching based on rectified views. However, these methods fail to explicitly exploit geometric relationships between multiple views, rendering them less capable of capturing the global dependencies, visibility, or scale changes. In this paper, we shift to a new perspective and propose a novel reference-free framework, dubbed FreeOmniMVS, via multi-view consistency maximization. The highlight of FreeOmniMVS is that it can aggregate pair-wise correlations into a robust, visibility-aware, and global consensus. As such, it is tolerant to occlusions, partial overlaps, and varying baselines. Specifically, to achieve global coherence, we introduce a novel View-pair Correlation Transformer (VCT) that explicitly models pairwise correlation volumes across all camera view pairs, allowing us to drop unreliable pairs caused by occlusion or out-of-focus observations. To realize scalable and visibility-aware consensus, we propose a lightweight attention mechanism that adaptively fuses the correlation vectors, eliminating the need for a designated reference view and allowing all cameras to contribute equally to the stereo matching process. Extensive experiments on diverse benchmark datasets demonstrate the superiority of our method for globally consistent, visibility-aware, and scale-aware omnidirectional depth estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。