arXiv:2606.10407cs.SDcs.CV2026-06

用目标检测方法精确定位热带雨林中鸟鸣的时频位置

Time-frequency localization of bird calls in dense soundscapes

论文配图:Time-frequency localization of bird calls in dense soundscapes
图 1 · 摘自论文原文
  • 将鸟鸣检测建模为谱图上的目标检测任务,使用YOLO11模型
  • 在新加坡数据上性能接近翻倍(IoMin@50 F1达81.8%)
  • 新标注工具与评估指标更适合处理声音边界模糊问题

被动声学监测可实现野生动物的大规模观测,但现有生物声学分类器通常仅预测物种在时间窗内的存在,无法精确定位发声事件在时间与频率上的位置,限制了后续分析。本文将鸟鸣检测视为谱图上的目标检测任务,训练YOLO11模型以在新加坡密集的热带声景中定位鸟鸣。同时,开发了一个开源浏览器标注工具,并提出交并比最小值(IoMin)作为评估指标,相比标准IoU更适用于处理模糊的声学边界。最佳YOLO模型在新加坡分布内数据上的表现近乎翻倍(IoMin@50 F1从42.1%提升至81.8%),且在未见的夏威夷分布外数据上仍优于基线(58.6% vs. 48.6%)。结果表明,目标检测框架是复杂声景中动物发声时频定位的有力方法。

原文摘要 · Abstract (English)

Passive acoustic monitoring enables large-scale observation of wildlife, but most bioacoustic classifiers only predict species presence in a time window without localizing vocalizations precisely in time or frequency, limiting downstream analyses. We formulate bird vocalization detection as an object detection task on spectrograms and train YOLO11 models to localize bird calls in dense tropical soundscapes from Singapore. We additionally introduce an open-source browser-based annotation tool and propose Intersection over Minimum (IoMin), an evaluation metric that better handles ambiguous acoustic boundaries than standard IoU and is better suited to the problem at hand. The best YOLO model nearly doubles baseline performance on in-distribution soundscapes from Singapore (81.8% vs. 42.1% IoMin@50 F1-score) while still outperforming the baseline on unseen out-of-distribution recordings from Hawaii (58.6% vs. 48.6%). These results suggest that object detection frameworks are a promising approach to time-frequency localization of animal vocalizations in complex soundscapes.

声景分析目标检测鸟鸣识别时频定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。