通过不确定性引导的数据增强,提升立体匹配模型在真实场景的泛化能力。
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
- 基于图像统计量扰动生成未见域数据,模拟多样化真实场景。
- 引入批级统计的高斯分布建模扰动不确定性,扩大训练域覆盖范围。
- 适合希望提升现有立体匹配模型跨域性能的研究者使用。
当前基于合成数据训练的先进立体匹配(SM)模型在面对真实数据域时泛化能力差,主要因色彩、光照、对比度和纹理等域间差异所致。本文提出一种不确定性引导的数据增强(UgDA)方法,认为RGB空间中的图像统计量(均值与标准差)承载了域特征,可通过合理扰动这些统计量生成未见域样本。为模拟更多潜在域,提出以批级统计为基础的高斯分布来建模扰动方向与强度的不确定性。此外,强制同一场景原始图与增强图间的特征一致性,促使模型学习结构感知且规避域依赖捷径的表示。该方法简单、与架构无关,可集成至任意SM网络。在多个挑战性基准上的大量实验表明,本方法显著提升了现有SM网络的泛化性能。
原文摘要 · Abstract (English)
State-of-the-art stereo matching (SM) models trained on synthetic data often fail to generalize to real data domains due to domain differences, such as color, illumination, contrast, and texture. To address this challenge, we leverage data augmentation to expand the training domain, encouraging the model to acquire robust cross-domain feature representations instead of domain-dependent shortcuts. This paper proposes an uncertainty-guided data augmentation (UgDA) method, which argues that the image statistics in RGB space (mean and standard deviation) carry the domain characteristics. Thus, samples in unseen domains can be generated by properly perturbing these statistics. Furthermore, to simulate more potential domains, Gaussian distributions founded on batch-level statistics are poposed to model the unceratinty of perturbation direction and intensity. Additionally, we further enforce feature consistency between original and augmented data for the same scene, encouraging the model to learn structure aware, shortcuts-invariant feature representations. Our approach is simple, architecture-agnostic, and can be integrated into any SM networks. Extensive experiments on several challenging benchmarks have demonstrated that our method can significantly improve the generalization performance of existing SM networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。