提出新方法精准检测重叠生物声学事件,适用于生态与保护研究。
Robust detection of overlapping bioacoustic sound events
- 基于起始点检测,结合自监督音频编码器,提升重叠声事件识别能力
- 在7个数据集及新采集的雀鸟录音上实现最先进性能
- 适合需要高精度声学事件分析的生态学与动物行为研究者
我们提出一种针对生物声学事件的鲁棒检测方法,可有效应对重叠事件这一常见问题,广泛应用于动物行为学、生态学和保护研究。不同于传统基于帧的多标签方法,本文提出名为Voxaboxen的起始点检测方法,借鉴计算机视觉中的目标检测思想,并融合近期自监督音频编码器的进展。该方法在每个时间窗口预测是否存在发声起始点及其持续时长,同时反向预测是否存在发声结束点及其开始时间。两组边界框通过图匹配算法融合。我们还发布了一个新数据集,用于评估重叠发声检测性能,包含标注精确的斑胸草雀录音,具有频繁重叠特征。在七个现有数据集及新数据集上测试表明,Voxaboxen优于自然基线和现有声事件检测方法,且在频繁重叠场景下性能依然稳定,表现最佳。
原文摘要 · Abstract (English)
We propose a method for accurately detecting bioacoustic sound events that is robust to overlapping events, a common issue in domains such as ethology, ecology and conservation. While standard methods employ a frame-based, multi-label approach, we introduce an onset-based detection method which we name Voxaboxen. It takes inspiration from object detection methods in computer vision, but simultaneously takes advantage of recent advances in self-supervised audio encoders. For each time window, Voxaboxen predicts whether it contains the start of a vocalization and how long the vocalization is. It also does the same in reverse, predicting whether each window contains the end of a vocalization, and how long ago it started. The two resulting sets of bounding boxes are then fused using a graph-matching algorithm. We also release a new dataset designed to measure performance on detecting overlapping vocalizations. This consists of recordings of zebra finches annotated with temporally-strong labels and showing frequent overlaps. We test Voxaboxen on seven existing data sets and on our new data set. We compare Voxaboxen to natural baselines and existing sound event detection methods and demonstrate SotA results. Further experiments show that improvements are robust to frequent vocalization overlap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。