收集5500+物种23万段音频,用于动物声音识别与生态研究。
The iNaturalist Sounds Dataset
- 基于全球公民科学平台数据,构建多物种音频集。
- 弱标签下仍可作为有效预训练资源,提升下游任务性能。
- 适合生态学、生物多样性研究及公众参与型应用开发。
我们提出 iNaturalist Sounds 数据集(iNatSounds),包含来自超过5,500个物种的230,000段音频,由全球27,000多名记录者贡献。数据涵盖鸟类、哺乳类、昆虫、爬行类和两栖类的声音,音频与物种标签均源自 iNaturalist 平台上的用户观测记录。每段录音长度不一,仅标注单一物种。我们对比了多种主干模型架构,评估多分类与多标签目标函数的表现。尽管标签为弱监督,但 iNatSounds 在强标签下游数据集上仍展现出良好的预训练效果。该数据集以单一免费存档形式开放,推动相关领域研究发展。我们期望基于此数据训练的模型能支持下一代公众参与应用,并助力生物学家、生态学家及土地管理者处理大规模音频数据,增进对多样化声景中物种构成的理解。
原文摘要 · Abstract (English)
We present the iNaturalist Sounds Dataset (iNatSounds), a collection of 230,000 audio files capturing sounds from over 5,500 species, contributed by more than 27,000 recordists worldwide. The dataset encompasses sounds from birds, mammals, insects, reptiles, and amphibians, with audio and species labels derived from observations submitted to iNaturalist, a global citizen science platform. Each recording in the dataset varies in length and includes a single species annotation. We benchmark multiple backbone architectures, comparing multiclass classification objectives with multilabel objectives. Despite weak labeling, we demonstrate that iNatSounds serves as a useful pretraining resource by benchmarking it on strongly labeled downstream evaluation datasets. The dataset is available as a single, freely accessible archive, promoting accessibility and research in this important domain. We envision models trained on this data powering next-generation public engagement applications, and assisting biologists, ecologists, and land use managers in processing large audio collections, thereby contributing to the understanding of species compositions in diverse soundscapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。