用语音模型处理动物叫声,跨物种效果好。
Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
- 用语音自监督模型提取动物叫声的深层特征
- 在多个物种上达到与专用模型相当的识别精度
- 适合对生物声学感兴趣的研究者快速入门
自监督语音模型在语音处理中表现优异,但在非语音数据上的应用尚未充分探索。本文研究了HuBERT、WavLM和XEUS等模型在生物声学检测与分类任务中的迁移能力。结果表明,这些模型能在不同物种间生成丰富的动物叫声潜在表示。通过线性探测分析时间平均表示的特性,并引入其他下游架构以捕捉时序信息。最后,探讨了频率范围和噪声对性能的影响。实验显示,该方法性能可媲美微调过的生物声学预训练模型,且噪声鲁棒的预训练策略显著提升效果。这些发现表明,基于语音的自监督学习为推进生物声学研究提供了一种高效框架。
原文摘要 · Abstract (English)
Self-supervised speech models have demonstrated impressive performance in speech processing, but their effectiveness on non-speech data remains underexplored. We study the transfer learning capabilities of such models on bioacoustic detection and classification tasks. We show that models such as HuBERT, WavLM, and XEUS can generate rich latent representations of animal sounds across taxa. We analyze the models properties with linear probing on time-averaged representations. We then extend the approach to account for the effect of time-wise information with other downstream architectures. Finally, we study the implication of frequency range and noise on performance. Notably, our results are competitive with fine-tuned bioacoustic pre-trained models and show the impact of noise-robust pre-training setups. These findings highlight the potential of speech-based self-supervised learning as an efficient framework for advancing bioacoustic research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。