arXiv:2410.23279cs.SDcs.AI2024-10中稿 · ASRU 2025被引 1

用自监督预训练提升猴子叫声识别的准确率与稳定性。

Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification

  • 用掩码自编码器在海量无标注录音上预训练变压器模型。
  • 在小样本噪声数据上实现比卷积网络更优的分段、分类与识别效果。
  • 适合研究非人灵长类语音行为或低资源语音分析的学者参考。

狨猴是一种高度发声的灵长类动物,是研究社交交流行为的重要模型。与人类语言不同,狨猴的叫声结构松散、变化大,且常在嘈杂、低资源条件下录制。学习其沟通模式需同时完成叫声分段、分类和发声者识别——这是一组具有挑战性的任务。以往的卷积神经网络虽能捕捉局部模式,但难以处理长时序依赖。我们采用使用自注意力机制的变换器来建模全局依赖,但在小规模、带噪声的标注数据集上出现过拟合和不稳定性。为此,我们在数百小时的未标注狨猴录音上,使用掩码自编码器(MAE)进行预训练。该预训练显著提升了模型的稳定性和泛化能力。结果表明,经MAE预训练的变换器优于卷积网络,证明现代自监督架构可有效建模低资源非人类语音通信。

原文摘要 · Abstract (English)

The marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, and recorded in noisy, low-resource conditions. Learning marmoset communication requires joint call segmentation, classification, and caller identification -- challenging domain tasks. Previous CNNs handle local patterns but struggle with long-range temporal structure. We applied Transformers using self-attention for global dependencies. However, Transformers show overfitting and instability on small, noisy annotated datasets. To address this, we pretrain Transformers with MAE -- a self-supervised method reconstructing masked segments from hundreds of hours of unannotated marmoset recordings. The pretraining improved stability and generalization. Results show MAE-pretrained Transformers outperform CNNs, demonstrating modern self-supervised architectures effectively model low-resource non-human vocal communication.

语音识别自监督学习灵长类语音变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。