arXiv:2606.02341cs.SDcs.LG2026-06

用双编码器融合波形与频谱,提升水下声学分类准确率。

Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification

论文配图:Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification
图 1 · 摘自论文原文
  • 双分支结构分别处理波形和频谱,结合可微模糊聚合机制。
  • 在DeepShip和ShipsEar数据集上准确率优于单编码器基线。
  • 参数高效微调,适合小样本水下声学任务,兼具可解释性。

水下声学分类广泛应用于海洋场景,但复杂声学环境带来挑战。传统方法多使用波形或频谱作为特征:频谱能建模谐波依赖,但可能过滤关键判别特征;波形包含相位信息,却因噪声大、复杂度高难以直接处理。本文提出一种双编码器神经架构,同时处理波形与频谱,采用预训练主干网络和参数高效微调模块,实现领域自适应。为融合两分支输出,引入基于可微Choquet积分的新型模糊聚合机制,动态平衡时序与频谱表示。该策略不仅提升分类精度,还提供可解释性——通过分析学习到的模糊测度,揭示类别特异性表示依赖变化。动态调整注意力以聚焦受非平稳信道失真影响较小的表示,缓解水下环境非平稳问题。在DeepShip和ShipsEar数据集上的实验表明,该架构相较独立单编码器基线取得显著性能提升,同时大幅减少可训练参数,降低过拟合风险与计算开销。

原文摘要 · Abstract (English)

Underwater acoustic classification has a wide array of oceanic applications, but faces challenges due to an increasingly complex acoustic environment. Waveform and spectrogram representations have been primarily used as acoustic data features for classification tasks in this domain. Spectrograms model harmonic dependencies, but these reduced representations can filter out acoustic features relevant for discrimination. While phase information from the waveform allows full characterization of the signal, the original waveform can be noisy and complex, rendering this representation difficult for models to process directly. This paper proposes a dual-encoder neural architecture to simultaneously process acoustic waveforms and spectrograms, leveraging pre-trained backbones and parameter-efficient fine-tuning modules, enabling a domain adaptation. To combine these adapted branches, a novel differentiable fuzzy aggregation mechanism based on the Choquet integral is introduced to balance the temporal and spectral representations. This fusion strategy not only yields higher classification accuracy but also provides interpretability. Specifically, by analyzing the learned fuzzy measures, insights are revealed about class-specific shifts in the network's representation reliance. By dynamically shifting attention to the representation least corrupted by potential asymmetric channel distortions, the proposed gating mechanism mitigates the non-stationary challenges of the underwater environment. Evaluations on the DeepShip and ShipsEar datasets demonstrate that the proposed architecture achieves classification improvements over independent single-encoder baselines, while simultaneously restricting the trainable parameter space. This mitigates the risk of overfitting on limited acoustic datasets while alleviating the computational costs associated with fully fine-tuning foundation models.

水下声学双编码器参数高效可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。