arXiv:2508.11845cs.SDcs.AI2025-08被引 9

提出通用动物发声编码器,提升多任务生物声学分析性能。

AVEX: What Matters for Animal Vocalization Encoding

  • 用自监督预训练+跨物种混合数据微调,提升编码效果。
  • 在26个数据集上实现当前最佳性能,涵盖物种识别等任务。
  • 强调数据多样性对模型泛化至关重要,适合生态研究者使用。

生物声学研究动植物发声,在保护、生物多样性监测和行为分析中具有重要意义。许多任务如物种、个体和行为分类与检测均适合机器学习,但常受限于标注数据稀缺,亟需能提取通用表征的生物声学编码器。现有编码器通常局限于鸟类等少数物种,依赖单一架构或训练方式,且评估范围窄。本文开展大规模实证研究,覆盖训练数据多样性与规模、模型架构与训练策略、评估任务与数据集广度等此前较少关注的方面。在26个数据集(包括物种分类、检测、个体识别、发声谱系发现等任务)上,通过自监督预训练后,在混合生物声学与通用音频数据上进行监督微调,取得当前最优的分布内与分布外性能。研究发现,两个阶段的数据多样性均至关重要。为支持持续研究,我们将公开模型检查点。

原文摘要 · Abstract (English)

Bioacoustics, the study of sounds produced by living organisms, plays a vital role in conservation, biodiversity monitoring, and behavioral studies. Many tasks in this field, such as species, individual, and behavior classification and detection, are well-suited to machine learning. However, they often suffer from limited annotated data, highlighting the need for a general-purpose bioacoustic encoder capable of extracting useful representations for diverse downstream tasks. Such encoders have been proposed before, but are often limited in scope due to a focus on a narrow range of species (typically birds), and a reliance on a single model architecture or training paradigm. Moreover, they are usually evaluated on a small set of tasks and datasets. In this work, we present a large-scale empirical study that covers aspects of bioacoustics that are relevant to research but have previously been scarcely considered: training data diversity and scale, model architectures and training recipes, and the breadth of evaluation tasks and datasets. We obtain encoders that are state-of-the-art on the existing and proposed benchmarks. We also identify what matters for training these encoders, such that this work can be extended when more data are available or better architectures are proposed. Specifically, across 26 datasets with tasks including species classification, detection, individual ID, and vocal repertoire discovery, we find self-supervised pre-training followed by supervised post-training on a mixed bioacoustics + general-audio corpus yields the strongest in- and out-of-distribution performance. We show the importance of data diversity in both stages. To support ongoing research and application, we will release the model checkpoints.

生物声学自监督学习编码器动物发声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。