对比多种迁移学习方法,提升鸟类声音分类的泛化能力。
Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
- 使用微调和知识蒸馏,跨数据集提升分类性能
- 浅层微调在复杂声景中表现更优,更具鲁棒性
- 建议完善标注细节,提升模型泛化能力
动物声音可由机器学习自动识别,在生物多样性监测中具有重要意义。尽管模型能力不断提升,但现有生物声学物种分类器在不同物种和生境间仍存在性能不平衡,尤其在复杂声景中表现不佳。本研究探讨了迁移学习在大规模鸟类声音分类中的有效性,涵盖单标签与多标签场景,以及CNN与Transformer等多种模型架构。实验表明,微调与知识蒸馏均能取得优异性能,其中交叉蒸馏在Xeno-canto数据上的域内表现尤为突出。然而,在泛化到真实声景时,浅层微调优于知识蒸馏,展现出更强的稳健性与约束性。研究还探讨了多物种标签不完整情况下的利用策略,呼吁动物声音领域加强标注实践,包括标注背景物种及提供时间信息,以训练更鲁棒的鸟类声音分类器。这些发现为预训练模型的最优复用提供了指导。
原文摘要 · Abstract (English)
Animal sounds can be recognised automatically by machine learning, and this has an important role to play in biodiversity monitoring. Yet despite increasingly impressive capabilities, bioacoustic species classifiers still exhibit imbalanced performance across species and habitats, especially in complex soundscapes. In this study, we explore the effectiveness of transfer learning in large-scale bird sound classification across various conditions, including single- and multi-label scenarios, and across different model architectures such as CNNs and Transformers. Our experiments demonstrate that both fine-tuning and knowledge distillation yield strong performance, with cross-distillation proving particularly effective in improving in-domain performance on Xeno-canto data. However, when generalizing to soundscapes, shallow fine-tuning exhibits superior performance compared to knowledge distillation, highlighting its robustness and constrained nature. Our study further investigates how to use multi-species labels, in cases where these are present but incomplete. We advocate for more comprehensive labeling practices within the animal sound community, including annotating background species and providing temporal details, to enhance the training of robust bird sound classifiers. These findings provide insights into the optimal reuse of pretrained models for advancing automatic bioacoustic recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。