arXiv:2507.11812cs.SDeess.AS2025-07

用多源数据融合与注意力机制,实时重建高精度水下声速场。

A Multimodal Data Fusion Attention-Empowered Generative Adversarial Network for Real Time 3D Underwater Sound Speed Field Construction

  • 融合多模态数据,通过注意力机制捕捉全局空间特征相关性。
  • 在真实数据集上误差低于0.3 m/s,RMSE比CNN方法降低近50%。
  • 适合海洋声学、水下通信与导航系统研究者参考。

声速剖面(SSPs)是决定声信号传播模式的关键水下参数,直接影响水下通信的能量效率和定位系统的准确性。传统获取方法如匹配场处理(MFP)、压缩感知(CS)和深度学习(DL)通常依赖现场声纳测量,对水下观测系统的部署要求严苛。为克服这一限制,实现无需现场数据采集的高精度声速场重建,本文提出一种新型多模态数据融合生成对抗网络,增强残差注意力模块(MDF-RAGAN)。该架构结合注意力机制有效捕捉全局空间特征相关性,同时利用残差模块提取由海表温度(SST)变化引起的深海声速分布微小扰动。在公开真实世界数据集上的实验结果表明,所提模型优于现有先进方法,估计误差低于0.3 m/s。具体而言,相较于卷积神经网络(CNN)和空间插值(SITP)方法,其均方根误差(RMSE)几乎降低一半;相比均值剖面法,RMSE减少65.8%。这些结果凸显了多源融合与跨模态注意力在提升声速剖面重建精度与鲁棒性方面的有效性。

原文摘要 · Abstract (English)

Sound speed profiles (SSPs) are crucial underwater parameters that determine the propagation patterns of acoustic signals, directly influencing the energy efficiency of underwater communication and the accuracy of positioning systems. Conventional techniques for obtaining SSPs, such as matched field processing (MFP), compressive sensing (CS), and deep learning (DL), typically depend on on-site sonar measurements, which impose stringent requirements on the deployment of underwater observation systems. To overcome this limitation and enable high-precision sound speed field reconstruction without the need for on-site underwater data collection, we propose a novel multimodal data-fusion generative adversarial network enhanced with residual attention blocks (MDF-RAGAN). This architecture integrates attention mechanisms to capture global spatial feature correlations effectively, while residual modules are employed to extract subtle perturbations in deep-ocean sound velocity distribution caused by sea surface temperature (SST) variations. Experimental results on a public real-world dataset demonstrate that the proposed model outperforms other state-of-the-art methods, achieving an estimation error of less than 0.3 m/s. Specifically, MDF-RAGAN reduces the root mean square error (RMSE) by nearly half compared to convolutional neural network (CNN) and spatial interpolation (SITP) methods, and attains a 65.8\% RMSE reduction relative to the mean profile method. These results highlight the effectiveness of multi-source fusion and cross-modal attention in enhancing the accuracy and robustness of sound speed profile reconstruction.

声速场生成对抗网络多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。