arXiv:2510.22119cs.CV2025-10被引 1

用单目深度先验增强立体匹配,提升复杂场景泛化能力

CogStereo: Neural Stereo Matching with Implicit Spatial Cognition Embedding

  • 用单目深度特征隐式注入空间认知,指导匹配优化
  • 在多个数据集上达到最优性能,尤其在遮挡和弱纹理区域表现优异
  • 适合追求跨域泛化能力的立体视觉研究者与工业应用

深度立体匹配虽在基准数据集上通过微调取得显著进展,但在零样本泛化方面仍不及其他视觉任务中的基础模型。本文提出 CogStereo 框架,针对遮挡、弱纹理等挑战性区域,无需依赖特定数据集先验。该方法利用单目深度特征作为先验,将隐式空间认知嵌入精炼过程,捕捉超越局部对应关系的整体场景理解,确保结构一致的视差估计。其双条件精炼机制结合像素级不确定性与认知引导特征,实现全局一致性修正。在 Scene Flow、KITTI、Middlebury、ETH3D、EuRoc 及真实世界数据上的大量实验表明,CogStereo 不仅达成当前最优结果,更在跨域泛化上表现卓越,推动立体视觉向认知驱动方向发展。

原文摘要 · Abstract (English)

Deep stereo matching has advanced significantly on benchmark datasets through fine-tuning but falls short of the zero-shot generalization seen in foundation models in other vision tasks. We introduce CogStereo, a novel framework that addresses challenging regions, such as occlusions or weak textures, without relying on dataset-specific priors. CogStereo embeds implicit spatial cognition into the refinement process by using monocular depth features as priors, capturing holistic scene understanding beyond local correspondences. This approach ensures structurally coherent disparity estimation, even in areas where geometry alone is inadequate. CogStereo employs a dual-conditional refinement mechanism that combines pixel-wise uncertainty with cognition-guided features for consistent global correction of mismatches. Extensive experiments on Scene Flow, KITTI, Middlebury, ETH3D, EuRoc, and real-world demonstrate that CogStereo not only achieves state-of-the-art results but also excels in cross-domain generalization, shifting stereo vision towards a cognition-driven approach.

立体匹配空间认知跨域泛化深度先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。