arXiv:2508.01277cs.SDcs.LG2025-08综述被引 13

对比多种生物声学大模型,找出最适合迁移使用的最佳选择。

Foundation Models for Bioacoustics -- a Comparative Review

  • 系统分析预训练数据与模型架构,评估跨任务迁移能力
  • Perch 2.0 在鸟鸣识别中表现最佳,自监督模型也胜过专用模型
  • 注意力探测法能更好挖掘模型潜力,适合音频任务迁移

自动化生物声学分析对生物多样性监测和保护至关重要,需要能够适应多种生物声学任务的先进深度学习模型。本文全面回顾了大规模预训练生物声学基础模型,并系统研究其在多个生物声学分类任务中的可迁移性。通过分析预训练数据来源与基准测试,我们梳理了生物声学表征学习的发展现状。进一步,我们回顾了生物声学基础模型,剖析其训练数据、预处理、增强策略、网络结构与训练范式。此外,我们在 BEANS 与 BirdSet 基准上对部分模型进行了广泛实证研究,评估其在线性探测与注意力探测下的泛化能力。实验结果表明:Perch 2.0 在 BirdSet 上得分最高(受限评估),并在 BEANS 上线性探测表现最强;BirdMAE 是基于探测策略的最佳模型,在 BirdSet 上领先,仅次于 BEATs$_{NLM}$(NatureLM-audio 的编码器);注意力探测有助于充分释放基于 Transformer 模型的性能;在注意力探测下,基于 AudioSet 自监督训练的通用音频模型优于许多专用鸟类声音模型。这些发现为从业者选择合适模型进行新任务迁移提供了重要指导。

原文摘要 · Abstract (English)

Automated bioacoustic analysis is essential for biodiversity monitoring and conservation, requiring advanced deep learning models that can adapt to diverse bioacoustic tasks. This article presents a comprehensive review of large-scale pretrained bioacoustic foundation models and systematically investigates their transferability across multiple bioacoustic classification tasks. We overview bioacoustic representation learning by analysing pretraining data sources and benchmarks. On this basis, we review bioacoustic foundation models, dissecting the models' training data, preprocessing, augmentations, architecture, and training paradigm. Additionally, we conduct an extensive empirical study of selected models on the BEANS and BirdSet benchmarks, evaluating generalisability under linear and attentive probing. Our experimental analysis reveals that Perch~2.0 achieves the highest BirdSet score (restricted evaluation) and the strongest linear probing result on BEANS, building on diverse multi-taxa supervised pretraining; that BirdMAE is the best model among probing-based strategies on BirdSet and second on BEANS after BEATs$_{NLM}$, the encoder of NatureLM-audio; that attentive probing is beneficial to extract the full performance of transformer-based models; and that general-purpose audio models trained with self-supervised learning on AudioSet outperform many specialised bird sound models on BEANS when evaluated with attentive probing. These findings provide valuable guidance for practitioners selecting appropriate models to adapt them to new bioacoustic classification tasks via probing.

生物声学大模型迁移学习音频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。