Perch 2.0拓展多物种训练,提升生物声学模型泛化能力。
Perch 2.0: The Bittern Lesson for Bioacoustics
- 采用自蒸馏与原型学习,联合优化分类与声源预测任务。
- 在BirdSet和BEANS上达当前最优,海洋迁移任务超越专用模型。
- 揭示细粒度物种分类是生物声学预训练的稳健任务。
Perch 是一个高效的生物声学预训练模型,此前仅在鸟类数据上监督训练,提供数千种发声物种的即用分类分数及强表征用于迁移学习。本次发布的 Perch 2.0 将训练范围扩展至大规模多物种数据集,采用自蒸馏方法,结合原型学习分类器与新的声源预测训练目标。Perch 2.0 在 BirdSet 与 BEANS 基准测试中达到当前最优性能,并在海洋生物声学迁移任务中优于专用海洋模型,尽管其海洋训练数据极少。论文提出假设:细粒度物种分类是生物声学预训练中尤为稳健的任务。
原文摘要 · Abstract (English)
Perch is a performant pre-trained model for bioacoustics. It was trained in supervised fashion, providing both off-the-shelf classification scores for thousands of vocalizing species as well as strong embeddings for transfer learning. In this new release, Perch 2.0, we expand from training exclusively on avian species to a large multi-taxa dataset. The model is trained with self-distillation using a prototype-learning classifier as well as a new source-prediction training criterion. Perch 2.0 obtains state-of-the-art performance on the BirdSet and BEANS benchmarks. It also outperforms specialized marine models on marine transfer learning tasks, despite having almost no marine training data. We present hypotheses as to why fine-grained species classification is a particularly robust pre-training task for bioacoustics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。