arXiv:2504.20776cs.SDcs.AI2025-04被引 1

构建欧洲鸣虫声学识别数据集,支持精准物种分类。

ECOSoundSet: a finely annotated dataset for the automated acoustic identification of Orthoptera and Cicadidae in North, Central and temperate Western Europe

  • 整合200种蝗虫与24种蝉的精细标注录音,覆盖西欧多国。
  • 含10,653段录音,其中强标注数据按8:1:1划分训练/验证/测试集。
  • 适合生态监测、声学识别及深度学习研究者使用。

当前用于自然声景中欧洲昆虫自动声学识别的工具范围有限。为使算法跨场景识别各物种微妙复杂的声学特征,亟需大规模且生态异质性强的声学数据集。本文提出ECOSoundSet(欧洲蝗虫与蝉声学数据集),包含10,653段来自北、中及温带西欧(安道尔、比利时、丹麦、法国本土及科西嘉岛、德国、爱尔兰、卢森堡、摩纳哥、荷兰、英国、瑞士)的200种蝗虫和24种蝉的录音(含亚种共217和26个分类单元),部分通过南法与加泰罗尼亚实地采集,部分来自欧洲多位昆虫学家贡献。数据集由粗标签录音(仅知目标物种存在)与细标注录音(明确每段声音的时间与频率范围)组成。我们还提供了强标注数据的训练/验证/测试集划分,比例约为0.8:0.1:0.1,便于深度学习模型的训练与评估。该数据集可作为现有在线录音资源的重要补充,助力欧洲蝗虫与蝉声学分类深度学习算法的发展。

原文摘要 · Abstract (English)

Currently available tools for the automated acoustic recognition of European insects in natural soundscapes are limited in scope. Large and ecologically heterogeneous acoustic datasets are currently needed for these algorithms to cross-contextually recognize the subtle and complex acoustic signatures produced by each species, thus making the availability of such datasets a key requisite for their development. Here we present ECOSoundSet (European Cicadidae and Orthoptera Sound dataSet), a dataset containing 10,653 recordings of 200 orthopteran and 24 cicada species (217 and 26 respective taxa when including subspecies) present in North, Central, and temperate Western Europe (Andorra, Belgium, Denmark, mainland France and Corsica, Germany, Ireland, Luxembourg, Monaco, Netherlands, United Kingdom, Switzerland), collected partly through targeted fieldwork in South France and Catalonia and partly through contributions from various European entomologists. The dataset is composed of a combination of coarsely labeled recordings, for which we can only infer the presence, at some point, of their target species (weak labeling), and finely annotated recordings, for which we know the specific time and frequency range of each insect sound present in the recording (strong labeling). We also provide a train/validation/test split of the strongly labeled recordings, with respective approximate proportions of 0.8, 0.1 and 0.1, in order to facilitate their incorporation in the training and evaluation of deep learning algorithms. This dataset could serve as a meaningful complement to recordings already available online for the training of deep learning algorithms for the acoustic classification of orthopterans and cicadas in North, Central, and temperate Western Europe.

声学识别昆虫监测数据集深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。