arXiv:2608.14824eess.AScs.LG2026-08

无需训练的最近中心分类法在少量语音样本下表现更优

A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification

  • 用固定预训练声学嵌入计算类中心,查询样本归于最近中心
  • 在少量标注样本时,该方法在EV数据集上超越有训练的分类器
  • 适合资源有限、标注样本少的动物叫声分类任务

我们在固定预训练声学嵌入基础上,对大象叫声分类进行了无参数的分段评估。不关注哪个嵌入在全量标注数据上训练出最佳分类器,而是考察最简单分类器在每类标注样本数变化时的表现。每类由支持集嵌入均值表示,查询样本按欧氏距离归属最近中心。在Perch(ver.1)、Perch(ver.2)和HuBERT(base, layer 2)嵌入及MFCC特征上,采用N-way k-shot设置,在与训练基线相同的交叉验证协议下进行评估。通过100次重采样支持集量化采样噪声。在较小的低资源EV数据集上,使用更强的Perch(ver.1)和Perch(ver.2)嵌入的中心分类器,从单样本/类起就超过全训练逻辑回归分类器,从双样本/类起超越更强的循环分类器。在强监督端到端基线所训练的缩减叫声类型集上,该分类器从少数样本开始即达到并超越其平均精度(mAP)。在更大的LDC数据集上,由于标注样本充足,训练基线始终占优。在每类5个样本时,使用最强嵌入Perch(ver.2)的中心分类器在EV数据集上达mAP 0.542,LDC数据集上达0.368。当标注样本少且固定嵌入已编码区分特征时,无参数最近中心分类法为更优选择。

原文摘要 · Abstract (English)

We present a parameter-free episodic evaluation of nearest-centroid classification for elephant vocalisations on fixed pretrained acoustic embeddings, across the Elephant Voices (EV) and Linguistic Data Consortium (LDC) datasets. Rather than asking which embedding yields the best classifier when trained on all available labelled data, we ask how the simplest classifier performs as labelled exemplars per class are varied. Each class is represented by the mean of its support-set embeddings, and each query is assigned to the nearest centroid under squared Euclidean distance. We evaluate this centroid classifier on the Perch (ver. 1), Perch (ver. 2), and HuBERT (base, layer 2) embeddings, together with mel frequency cepstral coefficient (MFCC) features, in an N-way k-shot manner under the same cross-validation protocol as the trained baselines. A bootstrap over 100 resampled support sets quantifies the sampling noise. On the smaller, low-resource EV dataset, the centroid classifier using the stronger Perch (ver. 1) and Perch (ver. 2) embeddings overtakes the fully-trained logistic regression classifier from a single exemplar per class and the stronger recurrent classifier from two. Over the reduced set of call types on which the strongly-supervised end-to-end baseline was trained, the centroid classifier matches and then surpasses that baseline in mean average precision (mAP), from a few exemplars per class. On the larger LDC dataset, where labelled exemplars are abundant, the trained baselines retain their advantage at every k considered. At five exemplars per class, the centroid classifier using the strongest embedding, Perch (ver. 2), attains a mAP of 0.542 on the EV dataset and 0.368 on the LDC dataset. Parameter-free nearest-centroid classification is the stronger choice when labelled exemplars are few and the fixed embedding already encodes the features that separate the call types.

少样本学习声音分类动物叫声无参数模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。