通过多任务与专家机制,提升水下声学目标识别的鲁棒性。
Advancing Robust Underwater Acoustic Target Recognition through Multi-task Learning and Multi-Gate Mixture-of-Experts
- 设计辅助任务估计目标尺寸,实现多任务协同学习。
- 在ShipsEar数据集上超越现有单任务模型,性能达最新水平。
- 适合水下目标识别、声学信号处理领域的研究者参考。
水下声学目标识别已成为水下声学领域的重要研究方向。然而,真实水下声信号记录的稀缺限制了数据驱动模型从复杂信号中学习稳健特征的能力,影响其实际应用中的稳定性。为此,本文提出M3框架(多任务、多门控、多专家),通过关注目标内在属性来增强模型捕捉稳健模式的能力。该框架设计一个以估计目标尺寸等属性为核心的辅助任务,与识别任务共享参数,实现多任务学习,使模型聚焦于跨任务共享信息,以正则化方式识别稳健目标特征,提升泛化能力。此外,M3引入多专家与多门控机制,为不同水下信号分配独立参数空间,实现对复杂信号模式的细粒度差异化处理。在ShipsEar水下舰船辐射噪声数据集上开展大量实验,结果表明M3显著优于当前最先进的单任务识别模型,达到新的性能基准。
原文摘要 · Abstract (English)
Underwater acoustic target recognition has emerged as a prominent research area within the field of underwater acoustics. However, the current availability of authentic underwater acoustic signal recordings remains limited, which hinders data-driven acoustic recognition models from learning robust patterns of targets from a limited set of intricate underwater signals, thereby compromising their stability in practical applications. To overcome these limitations, this study proposes a recognition framework called M3 (Multi-task, Multi-gate, Multi-expert) to enhance the model's ability to capture robust patterns by making it aware of the inherent properties of targets. In this framework, an auxiliary task that focuses on target properties, such as estimating target size, is designed. The auxiliary task then shares parameters with the recognition task to realize multi-task learning. This paradigm allows the model to concentrate on shared information across tasks and identify robust patterns of targets in a regularized manner, thereby enhancing the model's generalization ability. Moreover, M3 incorporates multi-expert and multi-gate mechanisms, allowing for the allocation of distinct parameter spaces to various underwater signals. This enables the model to process intricate signal patterns in a fine-grained and differentiated manner. To evaluate the effectiveness of M3, extensive experiments were implemented on the ShipsEar underwater ship-radiated noise dataset. The results substantiate that M3 has the ability to outperform the most advanced single-task recognition models, thereby achieving the state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。