用表面法向量提升立体匹配在真实场景的泛化能力
Geometry Reinforced Efficient Attention Tuning Equipped with Normals for Robust Stereo Matching

- 引入法向量作为几何先验,融合图像特征生成更鲁棒的上下文表征
- 在非朗伯区域增强抗干扰能力,使误差降低8.5%以上
- 轻量化注意力设计支持高分辨率推理,速度提升19.2%
尽管过去十年图像驱动的立体匹配取得显著进展,但合成到真实场景的零样本泛化仍是未解难题。该问题主要源于跨域差异及图像纹理固有的模糊性,尤其在遮挡、无纹理、重复和非朗伯(镜面/透明)区域。为此,本文提出GREATEN框架,利用表面法向量作为域不变、物体内在且具区分性的几何线索,弥补图像纹理的不足。框架包含三个核心组件:首先,门控上下文-几何融合(GCGF)模块自适应抑制不可靠的上下文线索,并融合过滤后的图像特征与法向量驱动的几何特征,构建域不变且具区分性的上下文-几何表征;其次,镜面-透明增强(STA)策略提升GCGF在非朗伯区域对误导性视觉线索的鲁棒性;第三,稀疏注意力设计保留了GREATStereo的细粒度全局特征提取能力,同时大幅降低计算开销,包括稀疏空间(SSA)、稀疏双匹配(SDMA)和简单体(SVA)注意力。仅在合成数据如SceneFlow上训练,GREATEN-IGEV在多个真实数据集上表现卓越:相较于FoundationStereo、Monster-Stereo和DEFOM-Stereo,ETH3D误差降低30%,非朗伯增强集误差降低8.5%,KITTI-2015误差降低14.1%。此外,GREATEN-IGEV比GREAT-IGEV快19.2%,支持在Middlebury上进行3K高分辨率推理,视差范围达768。
原文摘要 · Abstract (English)
Despite remarkable advances in image-driven stereo matching over the past decade, Synthetic-to-Realistic ZeroShot (Syn-to-Real) generalization remains an open challenge. This suboptimal generalization performance mainly stems from cross-domain shifts and ill-posed ambiguities inherent in image textures, particularly in occluded, textureless, repetitive, and non-Lambertian (specular/transparent) regions. To improve Synto-Real generalization, we propose GREATEN, a framework that incorporates surface normals as domain-invariant, object-intrinsic, and discriminative geometric cues to compensate for the limitations of image textures. The proposed framework consists of three key components. First, a Gated Contextual-Geometric Fusion (GCGF) module adaptively suppresses unreliable contextual cues in image features and fuses the filtered image features with normal-driven geometric features to construct domain-invariant and discriminative contextual-geometric representations. Second, a Specular-Transparent Augmentation (STA) strategy improves the robustness of GCGF against misleading visual cues in non-Lambertian regions. Third, sparse attention designs preserve the fine-grained global feature extraction capability of GREATStereo for handling occlusion and texture-related ambiguities while substantially reducing computational overhead, including Sparse Spatial (SSA), Sparse Dual-Matching (SDMA), and Simple Volume (SVA) attentions. Trained exclusively on synthetic data such as SceneFlow, GREATEN-IGEV achieves outstanding Syn-to-Real performance. Specifically, it reduces errors by 30% on ETH3D, 8.5% on the non-Lambertian Booster, and 14.1% on KITTI-2015, compared to FoundationStereo, Monster-Stereo, and DEFOM-Stereo, respectively. In addition, GREATEN-IGEV runs 19.2% faster than GREAT-IGEV and supports high-resolution (3K) inference on Middlebury with disparity ranges up to 768.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。