arXiv:2605.21143cs.SDcs.LG2026-05

用深度学习区分自然、人类和生物声音,提升生态声景分析的准确性。

CoarseSoundNet: Building a reliable model for ecological soundscape analysis

论文配图:CoarseSoundNet: Building a reliable model for ecological soundscape analysis
图 1 · 摘自论文原文
  • 构建CoarseSoundNet模型,直接在真实录音中分类三类声音。
  • 加入静音类和时长约束后,人类与自然声音识别率显著提升。
  • 可作为预处理工具,替代人工标注,适合生态监测研究者使用。

声景由生物声(biophony)、地质声(geophony)和人为声(anthropophony)三类声音构成。当前生态声景研究的关键问题是这些成分间的相互作用,尤其是生物声如何响应地质声与人为声。然而,现有分析工具难以准确分离这三类声音。尽管近年机器学习方法尝试实现自动化分析,但大多依赖特定任务或干净数据,难以推广至嘈杂的被动声学监测(PAM)录音。本文提出一种清晰可复现的建模框架,构建CoarseSoundNet——一个在真实PAM条件下训练的深度学习模型,用于粗粒度声景分类。系统评估了模型结构、额外训练类别、数据组成与评估策略的影响。结果表明:增加目标域相似的PAM数据能提升性能;引入显式静音类有助于优化;针对不同类别设置决策阈值与时长约束,尤其对人为声与地质声效果显著。错误分析显示,人为声因掩蔽效应识别困难,而静音与昆虫声常被误判为地质声或生物声。最后,生态案例研究证实,使用CoarseSoundNet预过滤录音后,获得的声学指数趋势与人工标注结果相当,支持其作为生态声学分析的有效预处理工具。

原文摘要 · Abstract (English)

A soundscape is composed of three types of sound: biophony (sounds made by animals), geophony (natural abiotic sounds) and anthropophony (sounds made by humans). A key research question in the field of soundscape ecology is how these components interact with each other, specifically how biophony responds to geophony and anthropophony. Nevertheless, as of today, there are not many analytical instruments that enable the distinct quantification of these elements. Recent machine learning (ML) approaches aim to support automated analysis but often rely on task-specific or clean data, limiting generalisation to noisy passive acoustic monitoring (PAM) recordings. This study presents a clear and reproducible structure to build ML models for coarse soundscape classification and introduces CoarseSoundNet, a deep learning model trained to distinguish biophony, geophony, and anthropophony under realistic PAM conditions. We systematically investigate model architectures, the influence of an additional training class, data composition, and evaluation strategies. Our findings suggest that model performance improves with additional PAM data, especially when similar to the target domain, and by introducing an explicit silence class during training. Class-specific decision thresholds and duration-based constraints further enhance performance, particularly for anthropophony and geophony. Error analyses exhibit challenges for anthropophony due to masking effects and confusions for silence and insect sounds for geophony and biophony. Finally, we conduct an ecological case study which shows that pre-filtering recordings with CoarseSoundNet yields acoustic index trends comparable to ground-truth filtering, supporting its use as an effective preprocessing tool for ecoacoustic analyses.

声景分析深度学习生态监测音频分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。