Perch 2.0模型在水下生物声学任务中表现优异,少样本迁移学习效果突出。
Perch 2.0 transfers 'whale' to underwater tasks
- 利用14597种生物声音数据预训练,提取通用声学特征。
- 在海洋哺乳动物音频任务上,少样本迁移学习性能优于多数现有模型。
- 适合缺乏标注数据的水下声学分类研究,尤其适用于新物种快速建模。
Perch 2.0 是一个在14,597种生物声音数据(包括鸟类、哺乳类、两栖类和昆虫)上预训练的监督式生物声学基础模型,在多个基准测试中达到当前最佳性能。尽管其训练数据几乎不包含海洋哺乳动物声音,我们仍通过少样本迁移学习评估其在海洋哺乳动物及水下音频任务上的表现。采用该模型生成的嵌入向量进行线性探测,并与Perch 1.0、SurfPerch、AVES-bio、BirdAVES、Birdnet V2.3等开源可迁移模型对比。结果表明,Perch 2.0的嵌入在大多数任务中均表现出一致高精度,显著优于其他嵌入模型,因此在仅有少量标注样本时,推荐使用该模型构建新的海洋哺乳动物分类器。
原文摘要 · Abstract (English)
Perch 2.0 is a supervised bioacoustics foundation model pretrained on 14,597 species, including birds, mammals, amphibians, and insects, and has state-of-the-art performance on multiple benchmarks. Given that Perch 2.0 includes almost no marine mammal audio or classes in the training data, we evaluate Perch 2.0 performance on marine mammal and underwater audio tasks through few-shot transfer learning. We perform linear probing with the embeddings generated from this foundation model and compare performance to other pretrained bioacoustics models. In particular, we compare Perch 2.0 with previous multispecies whale, Perch 1.0, SurfPerch, AVES-bio, BirdAVES, and Birdnet V2.3 models, which have open-source tools for transfer-learning and agile modeling. We show that the embeddings from the Perch 2.0 model have consistently high performance for few-shot transfer learning, generally outperforming alternative embedding models on the majority of tasks, and thus is recommended when developing new linear classifiers for marine mammal classification with few labeled examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。