用动物叫声预训练的模型在生物声学上不如语音预训练模型,且无需额外微调。
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
- 对比语音与动物叫声预训练模型在生物声学任务中的表现
- 动物叫声预训练仅带来微小提升,语音预训练已接近最优
- 语音模型微调后效果不一,说明其通用特征已足够强大
自监督学习(SSL)基础模型作为通用特征提取器,在多个任务中表现出强大能力。已有研究表明,基于人类语音预训练的模型在生物声学处理中具有高度可迁移性。本文探讨两个问题:(i) 直接在动物叫声上预训练的SSL模型是否显著优于语音预训练模型;(ii) 在自动语音识别(ASR)任务上微调语音预训练模型能否提升生物声学分类性能。我们在三个不同的生物声学数据集和两种任务上进行对比分析。结果表明,基于生物声学数据预训练仅带来微小性能提升,多数情况下与语音预训练模型表现相当。在ASR任务上微调后的结果参差不齐,说明SSL预训练所学得的通用表征已非常适配生物声学任务。这些发现凸显了语音预训练SSL模型在生物声学中的鲁棒性,暗示达到最优性能未必需要大量微调。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) foundation models have emerged as powerful, domain-agnostic, general-purpose feature extractors applicable to a wide range of tasks. Such models pre-trained on human speech have demonstrated high transferability for bioacoustic processing. This paper investigates (i) whether SSL models pre-trained directly on animal vocalizations offer a significant advantage over those pre-trained on speech, and (ii) whether fine-tuning speech-pretrained models on automatic speech recognition (ASR) tasks can enhance bioacoustic classification. We conduct a comparative analysis using three diverse bioacoustic datasets and two different bioacoustic tasks. Results indicate that pre-training on bioacoustic data provides only marginal improvements over speech-pretrained models, with comparable performance in most scenarios. Fine-tuning on ASR tasks yields mixed outcomes, suggesting that the general-purpose representations learned during SSL pre-training are already well-suited for bioacoustic tasks. These findings highlight the robustness of speech-pretrained SSL models for bioacoustics and imply that extensive fine-tuning may not be necessary for optimal performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。