融合多传感器数据,提升水下机器人在复杂环境中的感知与定位能力
Enhancing Situational Awareness in Underwater Robotics with Multi-modal Spatial Perception
- 采用相机、IMU和声学设备多模态融合,增强水下空间感知
- 在特隆赫姆峡湾实测中实现实时可靠状态估计与高质量3D重建
- 适用于复杂水下作业场景,尤其适合多摄像头配置的深海探测任务
自主水下航行器(AUV)和遥控潜水器(ROV)需要强大的空间感知能力,包括同时定位与地图构建(SLAM),以支持远程和自主任务。视觉系统虽能低成本获取丰富色彩与纹理信息,实现语义理解,但水下光照衰减、后向散射和低对比度常导致图像质量严重下降,使传统视觉SLAM失效。且多数方法依赖单目或双目输入,难以扩展至多摄像头配置。为此,本文提出融合相机、惯性测量单元(IMU)和声学设备的多模态感知方案,提升情境意识并实现鲁棒实时SLAM。结合几何与基于学习的方法及语义分析,在特隆赫姆峡湾多次实地部署中对工作级ROV采集的数据进行实验。结果表明,在视觉挑战环境下仍可实现可靠的状态估计与高质量3D重建。同时讨论了系统限制,如传感器标定问题和学习方法局限性,指出了未来需深入研究的方向。
原文摘要 · Abstract (English)
Autonomous Underwater Vehicles (AUVs) and Remotely Operated Vehicles (ROVs) demand robust spatial perception capabilities, including Simultaneous Localization and Mapping (SLAM), to support both remote and autonomous tasks. Vision-based systems have been integral to these advancements, capturing rich color and texture at low cost while enabling semantic scene understanding. However, underwater conditions -- such as light attenuation, backscatter, and low contrast -- often degrade image quality to the point where traditional vision-based SLAM pipelines fail. Moreover, these pipelines typically rely on monocular or stereo inputs, limiting their scalability to the multi-camera configurations common on many vehicles. To address these issues, we propose to leverage multi-modal sensing that fuses data from multiple sensors-including cameras, inertial measurement units (IMUs), and acoustic devices-to enhance situational awareness and enable robust, real-time SLAM. We explore both geometric and learning-based techniques along with semantic analysis, and conduct experiments on the data collected from a work-class ROV during several field deployments in the Trondheim Fjord. Through our experimental results, we demonstrate the feasibility of real-time reliable state estimation and high-quality 3D reconstructions in visually challenging underwater conditions. We also discuss system constraints and identify open research questions, such as sensor calibration, limitations with learning-based methods, that merit further exploration to advance large-scale underwater operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。