arXiv:2605.12534cs.SDcs.LG2026-05被引 2

针对动物叫声的噪声抑制模型,效率远超语音增强方法。

BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations

论文配图:BioSEN: A Bio-acoustic Signal Enhancement Network for Animal Vocalizations
图 1 · 摘自论文原文
  • 三模块设计:时频特征提取、谐波结构捕捉、能量自适应门控
  • 在三个数据集上性能媲美甚至超过主流语音增强模型
  • 计算量极低,适合野外生物多样性监测应用

现有音频增强研究多聚焦人类语音,而生物声学因录音噪声大、动物声音特性独特而较少被关注。为此,我们借鉴语音增强方法,构建了专为生物声学信号设计的BioSEN模型。该模型包含三个模块:多尺度双轴注意力单元用于时频特征提取,生物谐波多尺度增强单元用于捕捉声音谐波结构,能量自适应门控连接单元通过频率权重避免叫声被误判为噪声。在三个生物声学数据集上的测试表明,BioSEN性能达到或超过当前最优语音增强模型,且计算开销远低于同类方法。结果验证了其在生物声学音频增强中的有效性,展现出在生物多样性监测与保护中的应用潜力。

原文摘要 · Abstract (English)

Most work in audio enhancement targets human speech, while bioacoustics is less studied due to noisy recordings and the distinct traits of animal sounds. To fill this gap, we adapt speech enhancement methods and build BioSEN, a model made for bioacoustic signals. BioSEN has three modules: a multi-scale dual-axis attention unit for time-frequency feature extraction, a bio-harmonic multi-scale enhancement unit for capturing harmonic structures, and an energy-adaptive gating connection unit that uses frequency weights to keep vocalizations from being removed as noise. Tests on three bioacoustic datasets show that BioSEN matches or exceeds state-of-the-art speech enhancement models while using far less computation. These results show BioSEN's strength for bioacoustic audio enhancement and its promise for biodiversity monitoring and conservation.

生物声学音频增强动物叫声节能模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。