arXiv:2608.24325cs.AI2026-08

让声呐与光学融合,提升水下感知可靠性。

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

论文配图:SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception
图 1 · 摘自论文原文
  • 将声呐视为原生模态,设计专用编码器和物理感知增强。
  • 在声呐独用时达72.0%准确率,融合时超基线25.1个百分点。
  • 适用于水下机器人、海洋监测等复杂环境感知场景。

可靠的水下感知需依赖多模态互补:光学相机捕捉外观与语义,但受浊度影响迅速退化;声呐保留几何结构,却具有独特的距离-方位结构和声学伪影。现有多模态大模型主要基于光学编码器,难以建模声呐或动态利用其与光学的互补性。本文提出SonarLLM,一种将声呐作为原生感知模态的声呐-光学多模态大模型。它结合声呐专用编码器、模态特异性物理感知特征增强及可靠性感知的分层融合机制,实现声学结构与光学语义对齐,并随感知质量变化动态调整贡献权重。同时引入SonarBench基准,覆盖识别、计数、视觉问答和描述生成四项任务,支持声呐独用、光学独用和融合三种输入设置。通过固定场景与声呐观测、改变光学退化程度,实现对跨模态互补性的受控评估。SonarLLM在声呐独用的识别、计数和VQA任务上取得72.0%宏平均准确率,较最强基线提升34.4个百分点;融合模式下达68.7%,领先基线25.1个百分点。随着浊度增加,融合优于光学的增益从6.0提升至36.0,表明声呐在恶劣光学条件下的互补价值显著提升。结果表明,鲁棒异构感知不仅依赖加入声呐,更在于根据其感知特性进行建模与加权。

原文摘要 · Abstract (English)

Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonar-optical complementarity. We propose SonarLLM, a sonar-optical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasks: recognition, counting, visual question answering, and captioning; and, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics.

水下感知多模态声呐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。