构建大规模水下立体匹配数据集UWStereo,提升模型泛化能力。
UWStereo: A Large Synthetic Dataset for Underwater Stereo Matching
- 基于合成数据生成29,568对水下立体图像,含精确密集视差标注。
- 新方法通过跨域图像重建与跨视图注意力机制,显著提升泛化性能。
- 适用于水下视觉、机器人导航等需要鲁棒立体匹配的场景。
尽管立体匹配技术取得进展,但复杂水下环境的应用仍不充分,主要受限于水下图像的能见度降低、对比度下降等不利因素,以及难以获取真实训练标签——即在水下环境中同时拍摄图像并准确估计像素级深度信息。为推动水下立体匹配的发展,本文提出一个大规模合成数据集UWStereo,包含29,568对合成立体图像,对左视图提供稠密且精确的视差标注。数据集设计了四种不同水下场景,涵盖珊瑚、船只和机器人等多种物体,并引入相机模型、光照和环境效应的多样化变化。相比现有水下数据集,UWStereo在规模、多样性、标注质量及图像逼真度方面均更优。为验证其有效性,我们采用九种先进算法作为基准进行系统评估,结果表明当前模型仍难以跨域泛化。为此,我们提出一种新策略:在立体匹配训练前先学习跨域掩码图像重建,并集成跨视图注意力增强模块,以聚合长程内容信息,提升模型泛化能力。
原文摘要 · Abstract (English)
Despite recent advances in stereo matching, the extension to intricate underwater settings remains unexplored, primarily owing to: 1) the reduced visibility, low contrast, and other adverse effects of underwater images; 2) the difficulty in obtaining ground truth data for training deep learning models, i.e. simultaneously capturing an image and estimating its corresponding pixel-wise depth information in underwater environments. To enable further advance in underwater stereo matching, we introduce a large synthetic dataset called UWStereo. Our dataset includes 29,568 synthetic stereo image pairs with dense and accurate disparity annotations for left view. We design four distinct underwater scenes filled with diverse objects such as corals, ships and robots. We also induce additional variations in camera model, lighting, and environmental effects. In comparison with existing underwater datasets, UWStereo is superior in terms of scale, variation, annotation, and photo-realistic image quality. To substantiate the efficacy of the UWStereo dataset, we undertake a comprehensive evaluation compared with nine state-of-the-art algorithms as benchmarks. The results indicate that current models still struggle to generalize to new domains. Hence, we design a new strategy that learns to reconstruct cross domain masked images before stereo matching training and integrate a cross view attention enhancement module that aggregates long-range content information to enhance the generalization ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。