arXiv:2505.23964cs.SDcs.LG2025-05被引 1

用可学习滤波器自动识别船只声纹,跨距离识别准确率达96.6%

Acoustic Classification of Maritime Vessels using Learnable Filterbanks

  • 设计可训练的时频前端,动态聚焦关键频率成分
  • 在不同距离下实现96.63%准确率,比旧方法高12个百分点
  • 适合水下声学监测、海洋环境智能感知场景

基于声学特征可靠监测与识别海上船只,受不同录音场景影响较大。一个鲁棒的分类框架需具备跨多样声学环境和变源-传感器距离的泛化能力。为此,我们提出一种深度学习模型,通过可训练的谱前处理与时间特征编码器,学习自适应的Gabor滤波器组,能动态强调不同频率成分。在斯特雷特·乔治亚海峡的VTUAD水听器数据上训练,模型CATFISH在不同源-传感器距离下达到96.63%的测试准确率,超越先前基准超12个百分点。本文展示了模型结构,论证了架构选择,分析了学习到的Gabor滤波器,并对传感器数据融合与注意力池化进行了消融研究。

原文摘要 · Abstract (English)

Reliably monitoring and recognizing maritime vessels based on acoustic signatures is complicated by the variability of different recording scenarios. A robust classification framework must be able to generalize across diverse acoustic environments and variable source-sensor distances. To this end, we present a deep learning model with robust performance across different recording scenarios. Using a trainable spectral front-end and temporal feature encoder to learn a Gabor filterbank, the model can dynamically emphasize different frequency components. Trained on the VTUAD hydrophone recordings from the Strait of Georgia, our model, CATFISH, achieves a state-of-the-art 96.63 % percent test accuracy across varying source-sensor distances, surpassing the previous benchmark by over 12 percentage points. We present the model, justify our architectural choices, analyze the learned Gabor filters, and perform ablation studies on sensor data fusion and attention-based pooling.

声纹识别深度学习水下监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。