首个水下声呐-视觉配对数据集,助力多模态感知研究
A Sonar-Visual Dataset for Cross-Modal Underwater Robot Perception

- 构建76000+对声呐与视觉图像,覆盖17次潜水、6个海域
- 用单目相机基准提升7倍鱼检测[email protected],验证跨模态潜力
- 提供标注工具与数据管道,适合水下机器人多模态研究者
水下机器人通常结合摄像头与声呐进行感知,以利用视觉的丰富语义和声呐的稳定测距能力。然而,由于缺乏配对的声呐-视觉数据集,跨模态映射学习仍不充分。本文提出SOVIS,一个用于跨模态水下感知的声呐-视觉数据集。SOVIS包含在特隆赫姆峡湾六个地点进行的17次潜水采集的超过76,000对配对帧,并配备端到端的数据清洗与同步流程。我们还设计了交互式标注工具,加速配对数据标注。最后,基于少量标注数据完成跨模态鱼类检测原型任务,在[email protected]上相较单目相机基线提升7倍。SOVIS为推进跨模态水下感知研究迈出第一步,支持如从单目图像生成密集声呐图等新方向。
原文摘要 · Abstract (English)
Underwater robots typically use both cameras and sonar for perception to leverage the rich semantic details of vision and the robust range measurements of acoustics. However, learning to map between these modalities via cross-modal prediction remains underexplored due to limited sonar-visual paired datasets. We present SOVIS, a sonar-visual dataset for cross-modal underwater perception. SOVIS comprises over 76,000 paired frames collected across 17 dives at six sites in the Trondheimfjord, supported by an end-to-end pipeline that cleans and synchronizes the cross-modal sensor data. We also introduce an interactive annotation tool designed to accelerate the labeling process for this paired data. Finally, we demonstrate a proof-of-concept cross-modal fish detection task using a small subset of labeled data, achieving a 7x improvement in [email protected] over a monocular camera baseline. SOVIS serves as the first step toward advancing cross-modal underwater perception research, enabling research directions such as dense sonar prediction from monocular images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。