arXiv:2512.06793cs.CV2025-12AAAI被引 1

提出轻量级立体匹配网络,显著提升未知场景下的泛化能力。

Generalized Geometry Encoding Volume for Real-time Stereo Matching

  • 用深度感知特征提取通用结构先验,指导代价聚合
  • 动态融合先验信息,增强未知场景中的匹配关系
  • 实时推理下零样本泛化性能领先,适合真实驾驶场景

实时立体匹配方法通常关注特定域内的性能提升,却忽视了真实应用中泛化能力的重要性。现有立体基础模型虽借助单目基础模型提升泛化性,但推理延迟高。为此,本文提出广义几何编码体(GGEV),一种具备强泛化能力的实时立体匹配网络。首先提取编码域不变结构先验的深度感知特征,作为代价聚合的引导;随后引入深度感知动态代价聚合模块(DDCA),自适应地将这些先验融入每个视差假设中,有效增强未见场景中的脆弱匹配关系。两项设计均轻量且互补,构建出具备强泛化能力的几何编码体。实验表明,GGEV在零样本泛化能力上超越所有现有实时方法,并在KITTI 2012、KITTI 2015和ETH3D基准上达到最先进水平。

原文摘要 · Abstract (English)

Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular foundation models (MFMs) to improve generalization, but typically suffer from substantial inference latency. To address this trade-off, we propose Generalized Geometry Encoding Volume (GGEV), a novel real-time stereo matching network that achieves strong generalization. We first extract depth-aware features that encode domain-invariant structural priors as guidance for cost aggregation. Subsequently, we introduce a Depth-aware Dynamic Cost Aggregation (DDCA) module that adaptively incorporates these priors into each disparity hypothesis, effectively enhancing fragile matching relationships in unseen scenes. Both steps are lightweight and complementary, leading to the construction of a generalized geometry encoding volume with strong generalization capability. Experimental results demonstrate that our GGEV surpasses all existing real-time methods in zero-shot generalization capability, and achieves state-of-the-art performance on the KITTI 2012, KITTI 2015, and ETH3D benchmarks.

立体匹配实时推理泛化能力几何编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。