用预训练音频模型+线性探测,低成本实现水下船舶目标识别
Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
- 冻结预训练模型,仅用线性探测提取船舶特征
- 线性探测可抑制录音环境干扰,准确识别船型
- 适合数据少、标注难的水下声学监测场景
航运噪声加剧了海洋声污染,威胁生态系统。被动声学监测(PAM)系统生成大量水下录音,但人工分析效率低下,亟需自动化方法。当前水下声学目标识别(UATR)多依赖监督学习,受限于标注数据稀缺。本文首次对多种跨域预训练音频模型在UATR中的迁移效果进行实证对比。模型权重冻结,利用嵌入向量进行分类、聚类与相似性评估。结果表明,嵌入空间的几何结构主要受录音设备与环境影响。但简单线性探测能有效消除此类干扰,提取出船型特征。该方法以极低计算成本实现高效自动识别,大幅减少高质量标注船音数据的需求。
原文摘要 · Abstract (English)
Increasing levels of anthropogenic noise from ships contribute significantly to underwater sound pollution, posing risks to marine ecosystems. This makes monitoring crucial to understand and quantify the impact of the ship radiated noise. Passive Acoustic Monitoring (PAM) systems are widely deployed for this purpose, generating years of underwater recordings across diverse soundscapes. Manual analysis of such large-scale data is impractical, motivating the need for automated approaches based on machine learning. Recent advances in automatic Underwater Acoustic Target Recognition (UATR) have largely relied on supervised learning, which is constrained by the scarcity of labeled data. Transfer Learning (TL) offers a promising alternative to mitigate this limitation. In this work, we conduct the first empirical comparative study of transfer learning for UATR, evaluating multiple pretrained audio models originating from diverse audio domains. The pretrained model weights are frozen, and the resulting embeddings are analyzed through classification, clustering, and similarity-based evaluations. The analysis shows that the geometrical structure of the embedding space is largely dominated by recording-specific characteristics. However, a simple linear probe can effectively suppress this recording-specific information and isolate ship-type features from these embeddings. As a result, linear probing enables effective automatic UATR using pretrained audio models at low computational cost, significantly reducing the need for a large amounts of high-quality labeled ship recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。