arXiv:2512.11545cs.SDcs.AI2025-12被引 1

用梅尔频谱图构建图结构,提升水下目标识别准确率

Graph Embedding with Mel-spectrograms for Underwater Acoustic Target Recognition

  • 将梅尔频谱转为图数据,结合Transformer与图网络建模信号拓扑
  • 在两个基准数据集上性能接近顶尖方法,识别准确率显著优于传统模型
  • 适合水下声学、海洋工程领域研究者,可解释性强

水下声学目标识别(UATR)因船舶辐射噪声复杂性和海洋环境多变性而极具挑战。尽管深度学习已取得良好效果,但现有模型多隐含假设声学数据位于欧氏空间,这不适用于具有非平稳、非高斯、非线性特性的水下声信号。为此,本文提出UATR-GTransformer,一种非欧几里得深度学习模型,融合Transformer与图神经网络(GNN)。模型包含三个核心组件:梅尔分块模块、GTransformer模块和分类头。梅尔分块模块将梅尔频谱图划分为重叠块,GTransformer模块利用Transformer编码器捕捉分块间的相互信息,生成梅尔图嵌入;随后,GNN通过建模局部邻域关系增强嵌入,全连接网络进一步进行特征变换。基于两个常用基准数据集的实验结果表明,UATR-GTransformer性能与当前最优方法相当。此外,可解释性分析显示该模型能有效提取丰富的频域信息,展现出在海洋工程中的应用潜力。

原文摘要 · Abstract (English)

Underwater acoustic target recognition (UATR) is extremely challenging due to the complexity of ship-radiated noise and the variability of ocean environments. Although deep learning (DL) approaches have achieved promising results, most existing models implicitly assume that underwater acoustic data lie in a Euclidean space. This assumption, however, is unsuitable for the inherently complex topology of underwater acoustic signals, which exhibit non-stationary, non-Gaussian, and nonlinear characteristics. To overcome this limitation, this paper proposes the UATR-GTransformer, a non-Euclidean DL model that integrates Transformer architectures with graph neural networks (GNNs). The model comprises three key components: a Mel patchify block, a GTransformer block, and a classification head. The Mel patchify block partitions the Mel-spectrogram into overlapping patches, while the GTransformer block employs a Transformer Encoder to capture mutual information between split patches to generate Mel-graph embeddings. Subsequently, a GNN enhances these embeddings by modeling local neighborhood relationships, and a feed-forward network (FFN) further performs feature transformation. Experiments results based on two widely used benchmark datasets demonstrate that the UATR-GTransformer achieves performance competitive with state-of-the-art methods. In addition, interpretability analysis reveals that the proposed model effectively extracts rich frequency-domain information, highlighting its potential for applications in ocean engineering.

水下识别图神经网络梅尔频谱声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。