用多尺度注意力增强水下生物声学去噪与识别,提升复杂噪声中动物叫声的检测准确率。
Multi-Representation Attention Framework for Underwater Bioacoustic Denoising and Recognition
- 通过频谱图分割生成生物能量软掩码,引导注意力聚焦关键区域。
- 在真实海洋录音上实现92.3%的识别准确率,显著降低误报率。
- 适合用于大规模、跨环境的海洋生物多样性监测系统。
在圣劳伦斯河口自动监测海洋哺乳动物面临巨大挑战:叫声从低频呻吟到超声波点击跨度大,常相互重叠,并嵌入变化的人为与环境噪声中。我们提出一种多步骤、注意力引导的框架,先对频谱图进行分割以生成生物相关能量的软掩码,再将掩码与原始输入融合,实现多频段去噪分类。通过加拿大萨盖内圣劳伦斯海洋公园研究站的真实录音数据验证,基于分割的注意力与中层融合可有效提升信号区分度,减少误检,并在多种环境条件和信噪比下产生可靠表征。在分布外(OOD)测试中,尽管高容量基线模型性能下降,该框架仍保持稳定,简单融合机制(门控、拼接)亦表现良好。这表明其能学习可迁移表征而非过拟合特定变换,适用于大规模真实世界生物多样性监测。所有实验设置下,该框架均显著优于基线模型,在分布内与分布外数据上均有明显准确率提升。
原文摘要 · Abstract (English)
Automated monitoring of marine mammals in the St. Lawrence Estuary faces extreme challenges: calls span low-frequency moans to ultrasonic clicks, often overlap, and are embedded in variable anthropogenic and environmental noise. We introduce a multi-step, attention-guided framework that first segments spectrograms to generate soft masks of biologically relevant energy and then fuses these masks with the raw inputs for multi-band, denoised classification. Image and mask embeddings are integrated via mid-level fusion, enabling the model to focus on salient spectrogram regions while preserving global context. Using real-world recordings from the Saguenay St. Lawrence Marine Park Research Station in Canada, we demonstrate that segmentation-driven attention and mid-level fusion improve signal discrimination, reduce false positive detections, and produce reliable representations for operational marine mammal monitoring across diverse environmental conditions and signal-to-noise ratios. Beyond in-distribution evaluation, we further assess the generalization of Mask-Guided Classification (MGC) under distributional shifts by testing on spectrograms generated with alternative acoustic transformations. While high-capacity baseline models lose accuracy in this Out-of-distribution (OOD) setting, MGC maintains stable performance, with even simple fusion mechanisms (gated, concat) achieving comparable results across distributions. This robustness highlights the capacity of MGC to learn transferable representations rather than overfitting to a specific transformation, thereby reinforcing its suitability for large-scale, real-world biodiversity monitoring. We show that in all experimental settings, the MGC framework consistently outperforms baseline architectures, yielding substantial gains in accuracy on both in-distribution and OOD data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。