arXiv:2605.14963cs.CV2026-05

提出零样本全景立体匹配方法,解决全景图像匹配难题

H-OmniStereo: Zero-Shot Omnidirectional Stereo Matching with Heading-Aligned Normal Priors

论文配图:H-OmniStereo: Zero-Shot Omnidirectional Stereo Matching with Heading-Aligned Normal Priors
图 1 · 摘自论文原文
  • 构建280万对合成全景立体图像数据集,支持大规模训练
  • 引入航向对齐法向估计器,提升跨视角匹配准确性
  • 无需微调即可适配真实相机,适合自动驾驶等全向感知场景

顶底球面等距投影图像上的立体匹配为全向感知提供了有效框架,因其垂直对齐的视差线可使用由大规模数据集和单目先验驱动的先进透视立体架构。然而,此类方法性能受限于全景立体数据集稀缺以及透视单目先验在球面畸变下的退化。为此,我们提出H-OmniStereo,一种零样本全景立体匹配框架。首先,构建包含超过280万对顶底等距投影立体图像的高质量合成数据集以扩大训练规模。其次,提出一种在航向对齐坐标系中运行的等距投影单目法向估计器。该设计提供鲁棒于畸变且跨视角一致的几何先验,有助于建立可靠的对应关系,同时提升训练效率并兼容训练-测试视场不匹配问题。大量实验表明,该方法在域外数据集上优于现有方法,并能仅用一个模型成功推广至真实消费级相机设置。模型与数据集将开源于https://github.com/JIANG-CX/H-OmniStereo。

原文摘要 · Abstract (English)

Stereo matching on top-bottom equirectangular images provides an effective framework for full-surround perception, as vertically aligned epipolar lines enable the use of advanced perspective stereo architectures that are largely driven by large-scale datasets and monocular priors. However, the performance of such adaptations is severely limited by the scarcity of omnidirectional stereo datasets and the degradation of perspective monocular priors under spherical distortions. To address these challenges, we propose H-OmniStereo, a zero-shot omnidirectional stereo matching framework. First, we construct high-quality synthetic dataset comprising over 2.8 million top-bottom equirectangular stereo pairs to scale up training. Second, we introduce an equirectangular monocular normal estimator, specifically operating in a heading-aligned coordinate system. Beyond providing distortion-robust and cross-view-consistent geometric priors for establishing reliable correspondences in stereo matching, this design boosts training efficiency and accommodates train-test FoV mismatches. Extensive experiments show that our approach achieves higher accuracy than existing methods on out-of-domain datasets and successfully generalizes to real-world consumer camera setups using a single model. The model and dataset will be released at https://github.com/JIANG-CX/H-OmniStereo.

立体匹配全景视觉零样本生成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。