arXiv:2606.23702eess.AScs.AI2026-06

提出首个水下调制识别多场景评估基准与融合模型,提升实际部署鲁棒性。

Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift

论文配图:Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift
图 1 · 摘自论文原文
  • 分层融合2D时频图与1D高阶谱特征,用双向注意力对齐与自适应门控增强互补性。
  • 在真实海试数据上达91.14%~94.86%准确率,比最强基线高超15.71个百分点。
  • 首个统一评估框架覆盖分布外多种挑战,适合水下通信系统研发与测试人员。

调制识别系统依赖异构信号表征:2D信号图像模态(如时频图、循环平稳图)捕捉结构模式,1D统计特征(如高阶功率谱)提供互补线索。在分布偏移下,这些模态退化不均,导致鲁棒融合成为实际部署的核心挑战。现有研究还受限于缺乏统一评估协议,无法系统区分不同偏移类型。本文通过联合基准与模型研究,在水下声学调制识别任务中解决上述问题。UAMR-ShiftBench是首个涵盖分布内、低信噪比、未见环境、未见通信参数及实测海试评估的统一基准,包含两次南海海试(3月与11月)采集的两个独立真实数据集。SCP-TriCA 模型分层融合STFT、循环平稳图与P2/P4(二阶/四阶功率谱)模态:先通过双向交叉注意力对齐双2D模态,再以样本自适应选择门融合1D统计模态。在UAMR-ShiftBench上,该模型实现95.33%分布内准确率和74.59%模拟分布外平均准确率,优于最强基线5.12个百分点;在两海试子集上分别达到91.14%和94.86%,领先基线15.71和23.00个百分点。消融实验验证了模态互补性与分层融合设计的有效性。代码与模型已开源。

原文摘要 · Abstract (English)

Modulation recognition systems rely on heterogeneous signal representations. 2D signal-image modalities such as time-frequency and cyclostationary maps capture structural patterns, while 1D statistical descriptors such as higher-order power spectra encode complementary cues. Under distribution shift, these modalities degrade unevenly, making robust fusion a central challenge for practical deployment. Progress is further limited by the lack of a unified evaluation protocol that systematically separates different shift types. This paper addresses both challenges through a joint benchmark-and-model study in underwater acoustic modulation recognition. UAMR-ShiftBench is the first benchmark to jointly cover in-distribution, low-SNR, unseen-environment, unseen-communication-parameter, and measured sea-trial evaluation under a single matched protocol, with two independent real-world subsets collected during two sea-trial campaigns conducted in March and November in the South China Sea. SCP-TriCA fuses STFT, cyclostationary, and P2/P4 (second- and fourth-order power spectra) modalities hierarchically: the two 2D modalities are first aligned through bidirectional cross-attention, and the 1D statistical modality is then incorporated through a sample-adaptive selective gate. On UAMR-ShiftBench, SCP-TriCA achieves 95.33% in-distribution accuracy and 74.59% simulated OOD average, outperforming the strongest baseline by 5.12 percentage points, and reaches 91.14% and 94.86% on the two sea-trial subsets, exceeding the best baseline by 15.71 and 23.00 percentage points respectively. Ablation results confirm that the gains stem from modality complementarity and the hierarchical fusion design. Code and models are available at https://github.com/ronglaiqian/UAMR-ShiftBench.

水下通信调制识别分布外泛化多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。