arXiv:2601.18086cs.SDeess.AS2026-01

用语音大模型识别水下目标,效果超99%。

From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition

  • 复用语音特征提取流程,用语音大模型做声学编码器
  • 在真实数据集上达99%准确率,跨域测试仍超96.67%
  • 适合缺乏标注数据的水下声学任务研究者

水下声学目标识别(UATR)在海洋应用中至关重要,但受限于标注数据少和海洋环境复杂,仍具挑战。本文探讨核心问题:训练于海量语音数据的语音大模型(SLMs)能否有效迁移至水下声学任务?为此,提出UATR-SLM框架,复用语音特征处理流程,将SLM作为声学编码器,并添加轻量分类器。在DeepShip和ShipsEar基准测试中,UATR-SLM实现超过99%的域内准确率,对不同信号长度保持强鲁棒性,跨域评估最高达96.67%。结果表明SLMs在UATR中具备强大可迁移性,为利用语音基础模型解决水下声学问题提供了新范式。

原文摘要 · Abstract (English)

Underwater acoustic target recognition (UATR) plays a vital role in marine applications but remains challenging due to limited labeled data and the complexity of ocean environments. This paper explores a central question: can speech large models (SLMs), trained on massive human speech corpora, be effectively transferred to underwater acoustics? To investigate this, we propose UATR-SLM, a simple framework that reuses the speech feature pipeline, adapts the SLM as an acoustic encoder, and adds a lightweight classifier.Experiments on the DeepShip and ShipsEar benchmarks show that UATR-SLM achieves over 99% in-domain accuracy, maintains strong robustness across variable signal lengths, and reaches up to 96.67% accuracy in cross-domain evaluation. These results highlight the strong transferability of SLMs to UATR, establishing a promising paradigm for leveraging speech foundation models in underwater acoustics.

水下识别语音模型迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。