统一估计视差与表面法向量,提升复杂场景下的3D视觉精度
GeoStereo: A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

- 融合前馈匹配与扩散模型,利用几何先验联合预测
- 零样本下在KITTI、NYUv2等数据集上达到最优视差精度
- 适合处理低光、反光、透明等挑战性场景的3D重建任务
立体匹配与表面法向量估计是3D视觉的基础任务。现有前馈式立体方法在困难区域仍难以获得可靠结果,主要因缺乏强几何先验。本文提出GeoStereo,一种统一的立体几何估计框架,利用强大的扩散先验联合预测视差与表面法向量。具体地,GeoStereo将前馈立体匹配流水线与基于扩散的法向量估计分支耦合。为实现两任务有效交互,引入视差到法向量的初始化策略,并构建向左视图的形变条件以支持扩散过程。该耦合设计使扩散分支提供强结构先验,增强病态区域的视差估计;而前馈分支则为法向量预测提供可靠几何引导。大量实验表明,GeoStereo在低光、高反射和透明物体等挑战场景中表现稳健。零样本设置下,在KITTI和NYUv2等多个基准上实现视差估计的最高排名(Rank-1),并在iBims-1和ScanNet等真实室内基准上取得最佳法向量估计精度。
原文摘要 · Abstract (English)
Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in challenging regions, mainly due to the lack of strong geometric priors. In this paper, we propose $\textbf{GeoStereo}$, a unified stereo geometry estimation framework that leverages powerful diffusion priors to jointly predict disparity and surface normals. Specifically, GeoStereo couples a feed-forward stereo matching pipeline with a diffusion-based normal estimation branch. To enable effective interaction between the two tasks, we introduce a disparity to normal initialization strategy and construct a warp to left-view condition for the diffusion process. This coupled design allows the diffusion branch to provide strong structural priors that enhance disparity estimation in ill-posed regions, while the feed-forward branch offers reliable geometric guidance for accurate normal prediction. Extensive experiments show that GeoStereo performs reliably in challenging scenarios, including low-light environments, highly reflective surfaces, and transparent objects. Under zero-shot settings, it achieves Rank-1 disparity estimation on multiple benchmarks, including KITTI and NYUv2, and delivers the best normal estimation accuracy on many real indoor benchmarks, such as iBims-1 and ScanNet. Project page: https://qz-wei.github.io/GeoStereo.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。