融合单目深度模型与几何约束,实现无纹理、透明等难题下的鲁棒立体匹配。
Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail
- 双分支架构融合立体与单目上下文信息,提升匹配鲁棒性。
- 在镜面、透明等挑战场景下零样本性能超越现有方法。
- 仅用合成数据训练,却在多个基准上达到顶尖水平。
我们提出 Stereo Anywhere,一种新型立体匹配框架,通过将几何约束与单目深度视觉基础模型(VFMs)的鲁棒先验相结合,实现互补信息融合。采用双分支架构,创新设计成本体积融合机制,有效应对无纹理区域、遮挡及非朗伯表面等关键挑战。基于自建的光学错觉数据集 MonoTrap 及多基准评估,验证了仅用合成数据训练的模型在零样本泛化上达到当前最优,显著优于现有方案,在镜面与透明物体等复杂场景中表现尤为稳健。
原文摘要 · Abstract (English)
We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching with learned contextual cues. Following this design, our framework introduces novel cost volume fusion mechanisms that effectively handle critical challenges such as textureless regions, occlusions, and non-Lambertian surfaces. Through our novel optical illusion dataset, MonoTrap, and extensive evaluation across multiple benchmarks, we demonstrate that our synthetic-only trained model achieves state-of-the-art results in zero-shot generalization, significantly outperforming existing solutions while showing remarkable robustness to challenging cases such as mirrors and transparencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。