arXiv:2505.03071eess.AS2025-05被引 16

一小时快速构建生物声学识别模型,助力生态监测

The Search for Squawk: Agile Modeling in Bioacoustics

  • 用预训练鸟鸣嵌入减少数据需求,提升模型泛化能力
  • 通过索引音频搜索高效生成训练数据集,支持快速迭代
  • 适合生态学家快速应对新声音识别挑战,无需机器学习背景

被动式声学监测(PAM)在理解动物种群与生态系统健康方面展现出巨大潜力。然而,从数百万小时的音频中提取洞见需要开发专用识别器,这通常需要大量训练数据和机器学习专业知识。本文提出一种通用、可扩展且数据高效的系统,可在一小时内为新型生物声学问题构建识别器。系统包含三个核心组件:1)针对鸟鸣分类预训练的高泛化性声学嵌入,显著降低数据依赖;2)索引音频搜索机制,实现训练数据集的高效构建;3)嵌入预计算支持高效的主动学习循环,以最小延迟持续提升分类器性能。生态学家在三项新案例研究中应用该系统:通过未知声音分析珊瑚礁健康状况;识别夏威夷雏鸟叫声以量化繁殖成功率并改善濒危物种监测;以及圣诞岛鸟类分布建模。我们还通过模拟实验系统性探索设计选择,确立最佳实践。整体实验验证了系统的可扩展性、高效性与泛化能力,使科学家能快速应对新的生物声学挑战。

原文摘要 · Abstract (English)

Passive acoustic monitoring (PAM) has shown great promise in helping ecologists understand the health of animal populations and ecosystems. However, extracting insights from millions of hours of audio recordings requires the development of specialized recognizers. This is typically a challenging task, necessitating large amounts of training data and machine learning expertise. In this work, we introduce a general, scalable and data-efficient system for developing recognizers for novel bioacoustic problems in under an hour. Our system consists of several key components that tackle problems in previous bioacoustic workflows: 1) highly generalizable acoustic embeddings pre-trained for birdsong classification minimize data hunger; 2) indexed audio search allows the efficient creation of classifier training datasets, and 3) precomputation of embeddings enables an efficient active learning loop, improving classifier quality iteratively with minimal wait time. Ecologists employed our system in three novel case studies: analyzing coral reef health through unidentified sounds; identifying juvenile Hawaiian bird calls to quantify breeding success and improve endangered species monitoring; and Christmas Island bird occupancy modeling. We augment the case studies with simulated experiments which explore the range of design decisions in a structured way and help establish best practices. Altogether these experiments showcase our system's scalability, efficiency, and generalizability, enabling scientists to quickly address new bioacoustic challenges.

生物声学生态监测主动学习音频识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。