arXiv:2608.27795cs.CV2026-08

首个同步3D声呐与可见光影像的水下数据集,助力机器人精准感知。

uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception

论文配图:uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception
图 1 · 摘自论文原文
  • 融合3D多波束声呐与RGB图像,实现水下多模态数据同步采集。
  • 包含110个场景、95,834组同步观测,累计277.6分钟真实水域数据。
  • 适合水下机器人感知、跨模态学习与三维场景理解研究者使用。

自主水下机器人的可靠感知至关重要。然而,在光照差和散射严重的环境下,光学相机性能下降。前向探测(2D)声学传感器在这些条件下仍有效,但仅能测量距离和方位,无法确定俯仰角,导致回波点无法在三维空间中准确定位,阻碍了3D场景理解与精确目标检测。为此,我们提出 extbf{uScenes},一个包含同步3D多波束声呐点云与RGB图像的多模态水下数据集。该数据集包含110个场景、95,834组同步观测,覆盖277.6分钟的真实水域采集数据。uScenes为水下传感器融合、跨模态表示学习与3D场景理解奠定了基础。代码与数据集已开源:https://github.com/era-research-lab/uScenes。

原文摘要 · Abstract (English)

Robust perception is essential for the deployment of autonomous underwater robots. However, optical cameras become unreliable under poor illumination and backscatter. Forward looking (2D) acoustic sensors remain effective under these conditions, but they measure range and bearing while leaving elevation unresolved, creating an ambiguity that prevents individual sonar returns from being localized in three dimensional (3D) space. This complicates the sensor use for 3D scene understanding and precise object detection. We introduce \textbf{uScenes}, a multimodal underwater dataset containing synchronized 3D multibeam sonar point clouds and RGB imagery. The dataset contains 110 scenes and 95,834 synchronized observation, representing 277.6 minutes of data collected across multiple field sessions. uScenes establishes a foundation for underwater sensor fusion, cross modal representation learning and 3D scene understanding. Code and datasets are given at https://github.com/era-research-lab/uScenes.

水下感知多模态数据集3D声呐传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。