用少量标注数据实现复杂声景中鸟类叫声的精准检测
Semi-supervised classification of bird vocalizations
- 基于半监督学习,仅需每类11个标注样本即可训练
- 在315类110种鸟上达到0.701的F0.5分数,优于BirdNET
- 适用于长期连续监测,适合生态研究者使用
鸟类种群变化可反映生态系统整体变动,是重要的环境指示物种。结合机器学习与被动声学技术,可实现无需人工干预的长期连续监测。然而现有方法大多依赖大量专家标注数据,且难以分辨密集声景中的重叠叫声。本文提出一种半监督鸟类声学检测器,可在频率分离条件下识别时间重叠叫声,并仅需极少标注样本进行训练。模型在新加坡的社区采集开源数据与长期声景录音上进行训练与评估,测试集包含315类、110种鸟类,平均每类仅11个标注样本,取得0.701的均值F0.5得分。相比状态领先模型BirdNET,该方法在103种鸟的测试集上表现更优,且标注数据显著减少。进一步在144小时连续声景数据上验证,尽管新加坡声景复杂导致误报率高,仍证明在极小标注数据下实现高精度检测的可行性。
原文摘要 · Abstract (English)
Changes in bird populations can indicate broader changes in ecosystems, making birds one of the most important animal groups to monitor. Combining machine learning and passive acoustics enables continuous monitoring over extended periods without direct human involvement. However, most existing techniques require extensive expert-labeled datasets for training and cannot easily detect time-overlapping calls in busy soundscapes. We propose a semi-supervised acoustic bird detector designed to allow both the detection of time-overlapping calls (when separated in frequency) and the use of few labeled training samples. The classifier is trained and evaluated on a combination of community-recorded open-source data and long-duration soundscape recordings from Singapore. It achieves a mean F0.5 score of 0.701 across 315 classes from 110 bird species on a hold-out test set, with an average of 11 labeled training samples per class. It outperforms the state-of-the-art BirdNET classifier on a test set of 103 bird species despite significantly fewer labeled training samples. The detector is further tested on 144 microphone-hours of continuous soundscape data. The rich soundscape in Singapore makes suppression of false positives a challenge on raw, continuous data streams. Nevertheless, we demonstrate that achieving high precision in such environments with minimal labeled training data is possible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。