用深度模型自动识别大象叫声,助力保护与管理。
Learning to rumble: Automated elephant call classification, detection and endpointing using deep architectures
- 帧级检测结合序列到序列的音频变换器,实现精准叫声定位
- 最佳模型在5类叫声分类中AUC达0.957,7类子叫声分类达0.979
- 首次实现大象叫声细粒度分类,适合生态保护与智能监测
本文研究持续录音中大象叫声的自动检测、分离与分类问题,以支持保护工作与环境管理。不同于以往基于片段的检测方式,本研究采用帧级检测,隐含实现了叫声端点定位。实验使用两个标注数据集,分别包含亚洲象和非洲象叫声。评估多种浅层与深层分类模型,发现音频频谱变换器(AST)在序列到序列架构下表现最优,且通过预训练的迁移学习进一步提升性能与效率。此外,首次尝试基于标准分类体系的子叫声分类任务,验证了变换器架构在此任务上的领先优势。最优模型在帧级二分类中平均精度(AP)达0.962,在5类叫声分类和7类子叫声分类中的受试者工作特征曲线下面积(AUC)分别达到0.957和0.979,均刷新或达到新基准。结果表明,全自动大象叫声检测与子叫声分类系统已可实现,为种群行为与状态分析提供关键信息。
原文摘要 · Abstract (English)
We consider the problem of detecting, isolating and classifying elephant calls in continuously recorded audio. Such automatic call characterisation can assist conservation efforts and inform environmental management strategies. In contrast to previous work in which call detection was performed at a segment level, we perform call detection at a frame level which implicitly also allows call endpointing, the isolation of a call in a longer recording. For experimentation, we employ two annotated datasets, one containing Asian and the other African elephant vocalisations. We evaluate several shallow and deep classifier models, and show that the current best performance can be improved by using an audio spectrogram transformer (AST), a neural architecture which has not been used for this purpose before, and which we have configured in a novel sequence-to-sequence manner. We also show that using transfer learning by pre-training leads to further improvements both in terms of computational complexity and performance. Finally, we consider sub-call classification using an accepted taxonomy of call types, a task which has not previously been considered. We show that also in this case the transformer architectures provide the best performance. Our best classifiers achieve an average precision (AP) of 0.962 for framewise binary call classification, and an area under the receiver operating characteristic (AUC) of 0.957 and 0.979 for call classification with 5 classes and sub-call classification with 7 classes respectively. All of these represent either new benchmarks (sub-call classifications) or improvements on previously best systems. We conclude that a fully-automated elephant call detection and subcall classification system is within reach. Such a system would provide valuable information on the behaviour and state of elephant herds for the purposes of conservation and management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。