用合成音景数据训练,让鸟类声音检测模型更鲁棒且省人工。
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
- 通过合成背景音与目标叫声生成带动态标签的音景数据
- 模型在源叫声多样性大幅降低时仍保持高准确率
- 适合生态监测中数据标注成本高的场景
被动声学监测(PAM)分析常受限于标注训练数据所需大量人工工作。本研究提出一种合成数据框架,仅需少量原始素材即可生成大规模、丰富标注的训练数据,提升生物声学检测模型的鲁棒性。该框架通过组合干净背景噪声与孤立目标叫声(短耳鸮)合成真实音景,并在合成过程中自动生成如边界框等动态标签。经此数据微调的模型在真实音景中表现良好,即使源叫声多样性显著降低,性能依然稳定,表明模型学习到了泛化特征而未过拟合。结果证明,合成数据生成是从小样本数据训练鲁棒生物声学探测器的有效策略。该方法显著降低人工标注成本,突破计算生物声学中的关键瓶颈,增强生态评估能力。
原文摘要 · Abstract (English)
Passive Acoustic Monitoring (PAM) analysis is often hindered by the intensive manual effort needed to create labelled training data. This study introduces a synthetic data framework to generate large volumes of richly labelled training data from very limited source material, improving the robustness of bioacoustic detection models. Our framework synthesises realistic soundscapes by combining clean background noise with isolated target vocalisations (little owl), automatically generating dynamic labels like bounding boxes during synthesis. A model fine-tuned on this data generalised well to real-world soundscapes, with performance remaining high even when the diversity of source vocalisations was drastically reduced, indicating the model learned generalised features without overfitting. This demonstrates that synthetic data generation is a highly effective strategy for training robust bioacoustic detectors from small source datasets. The approach significantly reduces manual labelling effort, overcoming a key bottleneck in computational bioacoustics and enhancing ecological assessment capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。