构建首个具备强零样本泛化的立体匹配基础模型
FoundationStereo: Zero-Shot Stereo Matching

- 基于百万级合成数据与自清洗流程训练
- 零样本跨域精度超越现有方法,实现强鲁棒性
- 适合需要快速部署、无需微调的视觉系统开发者
深度立体匹配在特定数据集上通过领域微调已取得显著进展,但实现强零样本泛化——这是其他计算机视觉任务中基础模型的标志性能力——对立体匹配仍具挑战。本文提出FoundationStereo,一个面向立体深度估计的基础模型,旨在实现强大的零样本泛化能力。首先,构建包含100万对立体图像的大规模合成数据集,具备高多样性与高逼真度,并设计自动自清洗管道以去除模糊样本。随后,引入多项网络架构组件以提升可扩展性:包括侧向微调特征主干,利用视觉基础模型中的丰富单目先验缓解仿真到真实差距;以及长程上下文推理机制,用于有效过滤代价体。这些组件共同带来跨域优异的鲁棒性与准确性,树立了零样本立体深度估计的新标准。
原文摘要 · Abstract (English)
Tremendous progress has been made in deep stereo matching to excel on benchmark datasets through per-domain fine-tuning. However, achieving strong zero-shot generalization - a hallmark of foundation models in other computer vision tasks - remains challenging for stereo matching. We introduce FoundationStereo, a foundation model for stereo depth estimation designed to achieve strong zero-shot generalization. To this end, we first construct a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism, followed by an automatic self-curation pipeline to remove ambiguous samples. We then design a number of network architecture components to enhance scalability, including a side-tuning feature backbone that adapts rich monocular priors from vision foundation models to mitigate the sim-to-real gap, and long-range context reasoning for effective cost volume filtering. Together, these components lead to strong robustness and accuracy across domains, establishing a new standard in zero-shot stereo depth estimation. Project page: https://nvlabs.github.io/FoundationStereo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。