arXiv:2504.06710cs.LG2025-04被引 17

对比15种生物声学模型的特征提取能力,评估其在新物种识别中的泛化潜力。

Clustering and novel class recognition: evaluating bioacoustic deep learning feature extractors

  • 通过聚类与kNN分类分析15个模型的嵌入空间结构。
  • 发现不同模型在未见物种上的嵌入分离度差异显著。
  • 适合关注模型跨物种泛化能力的研究者参考。

在计算生物声学中,深度学习模型由特征提取器和分类器组成。特征提取器将输入音频片段生成向量表示(称为嵌入),可输入分类器进行判断。现有分类性能基准仅适用于训练数据中的已知物种,且难以比较训练于不同类群的模型。本文分析了15种涵盖多种架构、数据和训练范式的生物声学模型的特征提取器所生成的嵌入,通过聚类与kNN分类评估其嵌入空间结构,实现对特征提取器本身的独立比较。该方法有助于评估模型超越训练类别之外的适应性与泛化潜力。

原文摘要 · Abstract (English)

In computational bioacoustics, deep learning models are composed of feature extractors and classifiers. The feature extractors generate vector representations of the input sound segments, called embeddings, which can be input to a classifier. While benchmarking of classification scores provides insights into specific performance statistics, it is limited to species that are included in the models' training data. Furthermore, it makes it impossible to compare models trained on very different taxonomic groups. This paper aims to address this gap by analyzing the embeddings generated by the feature extractors of 15 bioacoustic models spanning a wide range of setups (model architectures, training data, training paradigms). We evaluated and compared different ways in which models structure embedding spaces through clustering and kNN classification, which allows us to focus our comparison on feature extractors independent of their classifiers. We believe that this approach lets us evaluate the adaptability and generalization potential of models going beyond the classes they were trained on.

生物声学特征提取嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。