用仿真数据训练水下立体深度网络,实现高精度实时水下定位与建图。
SurfSLAM: Sim-to-Real Underwater Stereo Reconstruction For Real-Time SLAM
- 通过仿真数据+自监督微调,让水下立体匹配模型从虚拟环境迁移到真实场景。
- 在包含24,000对图像的真实沉船勘测数据上,实现厘米级轨迹估计和密集三维重建。
- 适合水下机器人感知、海洋探测、沉船测绘等实际应用开发者参考。
水下机器人的定位与建图是核心感知能力。立体相机可低成本直接估计度量深度,但水下图像受光衰减、视觉伪影和动态光照影响,且缺乏纹理,导致地面训练的立体深度网络无法直接迁移。同时,缺乏真实世界水下立体数据集用于神经网络监督训练。这使得基于立体的同步定位与地图构建(SLAM)面临严重挑战。为此,我们提出一种新框架,利用仿真数据进行端到端训练,并通过自监督微调实现从模拟到现实的迁移。结合学习到的深度预测,开发了SurfSLAM——一种融合立体相机、IMU、气压计和多普勒速度计(DVL)的实时水下SLAM方法。最后,我们采集了一个具有挑战性的沉船勘测数据集,包含超过24,000对立体图像,以及高精度稠密摄影测量模型和参考轨迹用于评估。大量实验表明,该训练方法显著提升水下立体估计性能,实现了复杂沉船区域的准确轨迹估计与三维重建。
原文摘要 · Abstract (English)
Localization and mapping are core perceptual capabilities for underwater robots. Stereo cameras provide a low-cost means of directly estimating metric depth to support these tasks. However, despite recent advances in stereo depth estimation on land, computing depth from image pairs in underwater scenes remains challenging. In underwater environments, images are degraded by light attenuation, visual artifacts, and dynamic lighting conditions. Furthermore, real-world underwater scenes frequently lack rich texture useful for stereo depth estimation and 3D reconstruction. As a result, stereo estimation networks trained on in-air data cannot transfer directly to the underwater domain. In addition, there is a lack of real-world underwater stereo datasets for supervised training of neural networks. Poor underwater depth estimation is compounded in stereo-based Simultaneous Localization and Mapping (SLAM) algorithms, making it a fundamental challenge for underwater robot perception. To address these challenges, we propose a novel framework that enables sim-to-real training of underwater stereo disparity estimation networks using simulated data and self-supervised finetuning. We leverage our learned depth predictions to develop SurfSLAM, a novel framework for real-time underwater SLAM that fuses stereo cameras with IMU, barometric, and Doppler Velocity Log (DVL) measurements. Lastly, we collect a challenging real-world dataset of shipwreck surveys using an underwater robot. Our dataset features over 24,000 stereo pairs, along with high-quality, dense photogrammetry models and reference trajectories for evaluation. Through extensive experiments, we demonstrate the advantages of the proposed training approach on real-world data for improving stereo estimation in the underwater domain and for enabling accurate trajectory estimation and 3D reconstruction of complex shipwreck sites.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。