用单目图像生成立体匹配数据,提升真实场景下的零样本匹配性能
Boosting Zero-shot Stereo Matching using Large-scale Mixed Images Sources in the Real World
- 结合单目深度估计与扩散模型,从单视角图生成密集立体匹配数据
- 利用伪单目深度标签和动态不变损失,在标注稀疏时仍保持高精度
- 融合视觉基础模型提取通用特征,适合少标注与跨域场景
立体匹配方法依赖密集的像素级真值标签,而真实世界数据集的标注成本高昂。标注数据稀缺及合成与真实图像间的领域差异带来了显著挑战。本文提出新框架 BooSTer,利用视觉基础模型和大规模混合图像源(包括合成、真实及单视角图像)。首先,为充分挖掘单视角图像潜力,设计了一种结合单目深度估计与扩散模型的数据生成策略,从单视角图像生成密集立体匹配数据。其次,针对真实数据集标注稀疏问题,引入单目深度估计模型的知识,采用伪单目深度标签与动态尺度-平移不变损失提供额外监督。此外,通过视觉基础模型作为编码器提取鲁棒且可迁移的特征,显著提升准确率与泛化能力。在多个基准数据集上的大量实验表明,该方法在有限标注与领域偏移场景下均显著优于现有方法。
原文摘要 · Abstract (English)
Stereo matching methods rely on dense pixel-wise ground truth labels, which are laborious to obtain, especially for real-world datasets. The scarcity of labeled data and domain gaps between synthetic and real-world images also pose notable challenges. In this paper, we propose a novel framework, \textbf{BooSTer}, that leverages both vision foundation models and large-scale mixed image sources, including synthetic, real, and single-view images. First, to fully unleash the potential of large-scale single-view images, we design a data generation strategy combining monocular depth estimation and diffusion models to generate dense stereo matching data from single-view images. Second, to tackle sparse labels in real-world datasets, we transfer knowledge from monocular depth estimation models, using pseudo-mono depth labels and a dynamic scale- and shift-invariant loss for additional supervision. Furthermore, we incorporate vision foundation model as an encoder to extract robust and transferable features, boosting accuracy and generalization. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach, achieving significant improvements in accuracy over existing methods, particularly in scenarios with limited labeled data and domain shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。