用时序卷积网络实现无需逐声波标注的生物声学分类
Classifying bioacoustic data without individual call annotations using temporal convolutional networks and feature extractors
- 基于时序卷积网络与特征提取器,直接从弱标注数据中学习
- 在4分钟录音上实现召回率0.83、误报率0.13,接近专家一致水平
- 自动特征提取比人工选特征更稳定,适合多场景应用
被动声学监测(PAM)生成的大规模生物声学数据常因标注困难而仅提供弱标签(如物种存在/不存在)。为有效捕捉长音频段中的复杂时序模式与关键特征,本文提出包含数据标准化、特征提取和时序卷积网络(TCN)分类的框架,无需设定启发式规则或耗时的强标签。以不同来源和部署条件下采集的抹香鲸(Physeter macrocephalus)click序列作为案例,在4分钟录音上,该方法达到超过0.83的召回率和0.13的假阳性率,性能接近专家间一致性水平。对比变分自编码器(VAEs)与传统人工特征选择两种提取方式,两者表现相近,但基于VAE的方法在跨数据集和录音条件下更具稳定性。结果表明,该框架可有效利用已有弱标注数据训练自动分类模型,突破以往弱标签带来的局限。
原文摘要 · Abstract (English)
Bioacoustic data from Passive Acoustic Monitoring (PAM) generates large datasets where obtaining detailed auditing and labelling is often impractical, resulting in weak annotations (e.g., presence/absence of species over several minutes of recording). In order to effectively capture the complex temporal patterns and key features of long audio segments, we propose a framework comprising dataset standardisation, feature extraction, and classification via Temporal Convolutional Networks (TCN). This approach eliminates the necessity for setting heuristic decision rules or creating time-consuming strong labels. To demonstrate the effectiveness of our approach, we use sperm whale (\textit{Physeter macrocephalus}) click trains in 4-minute recordings as a case study, from a dataset comprising diverse sources and deployment conditions to maximise generalisability. Our TCN classifiers achieve recall rates exceeding 0.83 at a 0.13 false positive rate, comparable to agreement rates between expert annotators. We compare two methods of feature extraction, Variational AutoEncoders (VAEs) and traditional handpicking of features, and found them to yield similar performance results, with the VAE-based classifiers seeing a more stable performance across datasets and recording conditions. These results offer a way forward in leveraging numerous existing annotated bioacoustic datasets to train automatic classification models, effectively overcoming previous limitations associated with weak labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。