arXiv:2503.02422cs.SDcs.LG2025-03被引 1

用不确定度聚合策略,让标注效率提升,少标8%标签也能达到全标注效果。

Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning

  • 按音频中最不确定片段选样本,不平均处理所有段落。
  • 仅用8%标签就达到全标注模型性能,适合声音事件稀疏场景。
  • 适合生物声学、时间序列数据的高效标注,尤其在资源有限时。

声事件检测(SED)应用中海量音频数据亟需高效标注策略以支持监督学习。人工标注成本高、耗时长,主动学习(AL)成为降低标注负担的可行方案。本文提出一种名为Top K Entropy的新颖不确定性聚合策略,不再对音频中所有片段的不确定性取均值,而是聚焦于最不确定的片段,从而优先选择整段音频进行标注,显著提升稀疏数据场景下的效率。我们在包含獴、狗叫和婴儿哭声等真实生物声学监测场景的公园音频混合数据集上进行了评估。结果表明,使用Top K Entropy进行主动学习,仅需8%的标注量即可实现与全标注数据训练相同模型性能;该方法优于均值熵(Mean Entropy),证明应由最不确定片段代表整段音频的不确定性。研究凸显了主动学习在音频及时间序列可扩展标注中的潜力。

原文摘要 · Abstract (English)

The vast amounts of audio data collected in Sound Event Detection (SED) applications require efficient annotation strategies to enable supervised learning. Manual labeling is expensive and time-consuming, making Active Learning (AL) a promising approach for reducing annotation effort. We introduce Top K Entropy, a novel uncertainty aggregation strategy for AL that prioritizes the most uncertain segments within an audio recording, instead of averaging uncertainty across all segments. This approach enables the selection of entire recordings for annotation, improving efficiency in sparse data scenarios. We compare Top K Entropy to random sampling and Mean Entropy, and show that fewer labels can lead to the same model performance, particularly in datasets with sparse sound events. Evaluations are conducted on audio mixtures of sound recordings from parks with meerkat, dog, and baby crying sound events, representing real-world bioacoustic monitoring scenarios. Using Top K Entropy for active learning, we can achieve comparable performance to training on the fully labeled dataset with only 8% of the labels. Top K Entropy outperforms Mean Entropy, suggesting that it is best to let the most uncertain segments represent the uncertainty of an audio file. The findings highlight the potential of AL for scalable annotation in audio and time-series applications, including bioacoustics.

主动学习生物声学音频标注稀疏数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。