arXiv:2409.12634cs.SDcs.AI2024-09被引 5

用人类语音训练的模型能更好区分蝙蝠叫声类型

Exploring bat song syllable representations in self-supervised audio encoders

  • 用人类语音预训练的音频编码器更擅长捕捉蝙蝠叫声差异
  • 不同音节类型在编码空间中区分度最高达87.3%
  • 为跨物种声学分析提供新思路,适合生态与语音研究者

深度学习模型在人类语音上预训练后,能否区分其他物种的发声类型?我们分析了多个自监督音频编码器对蝙蝠叫声音节的表征能力,发现基于人类语音预训练的模型能生成最具区分性的音节表征。该结果标志着跨物种迁移学习在蝙蝠生物声学中的初步应用,也深化了对音频编码器处理分布外信号的理解。

原文摘要 · Abstract (English)

How well can deep learning models trained on human-generated sounds distinguish between another species' vocalization types? We analyze the encoding of bat song syllables in several self-supervised audio encoders, and find that models pre-trained on human speech generate the most distinctive representations of different syllable types. These findings form first steps towards the application of cross-species transfer learning in bat bioacoustics, as well as an improved understanding of out-of-distribution signal processing in audio encoder models.

音频编码蝙蝠声学跨物种自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。