自监督模型ARIONet通过对比学习与未来帧预测,提升鸟类叫声分类准确率。
ARIONet: An Advanced Self-supervised Contrastive Representation Network for Birdsong Classification and Future Frame Prediction
- 基于Transformer的自监督网络,融合多维声学特征进行对比学习。
- 在4个数据集上分类准确率达98.41%,未来帧预测相似度达95%。
- 适合生态监测、生物多样性研究等无需大量标注的场景。
自动化鸟类叫声分类对推动生态监测和生物多样性研究至关重要。尽管近期已有进展,但现有方法仍严重依赖标注数据,特征表示有限,并忽视了物种识别所需的时序动态。本文提出一种自监督对比学习网络ARIONet(Acoustic Representation for Interframe Objective Network),通过增强音频表示联合优化对比分类与未来帧预测。模型在基于Transformer的编码器中整合多种互补声学特征。框架设计两个核心目标:(1) 通过最大化同一音频片段不同增强视图间的相似性并拉远不同样本的距离,学习具有区分性的物种特异性表示;(2) 通过预测未来音频帧建模时序动态,全程无需大规模标注。我们在四个多样化的鸟类叫声数据集上验证,包括英国鸟类叫声数据集、鸟类叫声数据集及两个扩展的Xeno-Canto子集(A-M 和 N-Z)。该方法持续优于现有基线,在四组数据上的分类准确率分别为98.41%、93.07%、91.89%和91.58%,F1分数分别为97.84%、94.10%、91.29%和90.94%。此外,未来帧预测任务中误差低,余弦相似度最高达95%。大量实验进一步证实该自监督策略在捕捉复杂声学模式与时序依赖方面的有效性,且具备实际生态保育与监测应用潜力。
原文摘要 · Abstract (English)
Automated birdsong classification is essential for advancing ecological monitoring and biodiversity studies. Despite recent progress, existing methods often depend heavily on labeled data, use limited feature representations, and overlook temporal dynamics essential for accurate species identification. In this work, we propose a self-supervised contrastive network, ARIONet (Acoustic Representation for Interframe Objective Network), that jointly optimizes contrastive classification and future frame prediction using augmented audio representations. The model simultaneously integrates multiple complementary audio features within a transformer-based encoder model. Our framework is designed with two key objectives: (1) to learn discriminative species-specific representations for contrastive learning through maximizing similarity between augmented views of the same audio segment while pushing apart different samples, and (2) to model temporal dynamics by predicting future audio frames, both without requiring large-scale annotations. We validate our framework on four diverse birdsong datasets, including the British Birdsong Dataset, Bird Song Dataset, and two extended Xeno-Canto subsets (A-M and N-Z). Our method consistently outperforms existing baselines and achieves classification accuracies of 98.41%, 93.07%, 91.89%, and 91.58%, and F1-scores of 97.84%, 94.10%, 91.29%, and 90.94%, respectively. Furthermore, it demonstrates low mean absolute errors and high cosine similarity, up to 95\%, in future frame prediction tasks. Extensive experiments further confirm the effectiveness of our self-supervised learning strategy in capturing complex acoustic patterns and temporal dependencies, as well as its potential for real-world applicability in ecological conservation and monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。