arXiv:2609.04411cs.ROcs.CV2026-09

用声呐监督单目水下图像,实现高精度鸟瞰图占位预测

AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision

论文配图:AquaBEV: Monocular Underwater BEV Occupancy with 3D Sonar Supervision
图 1 · 摘自论文原文
  • 通过声呐数据监督单目图像,构建无标定极坐标特征映射
  • 在水下基准上达到31.4%可见区域交并比,优于基线4.0%
  • 适合水下机器人导航与环境感知研究者使用

自主水下机器人广泛用于勘探、监测和检查,其安全导航依赖于对周围自由与占用空间的理解。鸟瞰图(BEV)占位提供了此类表征,但仅凭单张水下RGB图像预测存在困难,因外观提供的几何线索有限且不可靠。3D成像声呐可提供互补的几何测量以监督该任务。本文提出AquaBEV,一种从单张RGB图像预测局部BEV占位的模型,训练时利用配对的3D成像声呐作为几何监督。AquaBEV将视觉特征映射至无标定极坐标表示,并沿距离维度进行因果解码,最后重建为笛卡尔坐标系下的BEV预测。建立了一个受控的水下占位基准,统一协议下适配代表性方法完成对比。AquaBEV在可见区域获得31.4%的交并比,在观测区域达38.6%,相较最强迁移基线分别提升4.0%和4.3%。

原文摘要 · Abstract (English)

Autonomous underwater robots are widely used for exploration, monitoring, and inspection, where safe navigation depends on understanding the surrounding free and occupied space. Bird's eye view (BEV) occupancy provides such a representation, but predicting it from a single underwater RGB image is difficult due to limited, unreliable geometric cues from appearance alone. 3D imaging sonar offers complementary geometric measurements to supervise this task. We introduce AquaBEV, a monocular underwater occupancy model that predicts local BEV occupancy from a single RGB image, using paired 3D imaging sonar as geometric supervision during training. AquaBEV maps visual features into a calibration free polar representation and applies causal decoding along the range dimension before reconstructing the prediction in Cartesian BEV coordinates. A controlled underwater occupancy benchmark was established, adapting representative occupancy methods to the same RGB to sonar task under a unified protocol. AquaBEV achieves 31.4 Visible IoU and 38.6 Observed IoU, 4.0% and 4.3% relative improvements over the strongest transferred baseline.

水下感知鸟瞰图声呐监督单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。