arXiv:2606.12503cs.LGcs.SD2026-06被引 3

用五年的海豚鸣叫数据训练出首个专用自监督模型,可精细解析海豚交流模式。

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

论文配图:Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations
图 1 · 摘自论文原文
  • 基于海豚长期鸣叫数据,定制化改进Wav2Vec2.0架构
  • 在哨音分类与检测任务上超越通用模型表现
  • 生成的嵌入可识别鸣叫单元结构,适合动物通信研究

自监督学习(SSL)为生物声学带来新机遇,使无需人工标注即可大规模建模动物鸣叫成为可能。然而,当前该领域模型多注重跨物种泛化,难以揭示个体通信系统的细微结构。本文收集并发布超过五年、来自五只已知海豚的纵向录音数据,为研究海豚交流提供前所未有的资源。我们基于Wav2Vec2.0架构,提出首个专用于海豚的大型自监督模型Dolph2Vec,仅在此数据上训练。在签名哨音分类与哨音检测两个生物学相关任务上,Dolph2Vec显著优于通用基线模型。更重要的是,其学习到的嵌入与码本结构能捕捉与海豚哨音类别对齐的可解释声学单元,甚至可能反映子哨音结构,支持对交流模式的细粒度分析。结果表明,SSL不仅是建模工具,更可作为科学探索手段,助力动物交流研究中的假设检验。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has opened new opportunities in bioacoustics by enabling scalable modeling of animal vocalizations without the need for expensive manual annotation. However, current SSL models in this domain prioritize broad generalization across species and are not optimized for uncovering the fine-grained structure of individual communication systems. In this work, we collect and release a novel dataset of over five years of longitudinal recordings, from five known dolphins in a semi-naturalistic marine environment, an unprecedented resource for studying dolphin communication. We adapt the Wav2Vec2.0 Baevski et al. (2020) architecture to this domain and introduce Dolph2Vec, the first large-scale, species-specific SSL model trained exclusively on this data. We benchmark our model on two biologically relevant tasks: signature whistle classification and whistle detection. Dolph2Vec significantly outperforms general-purpose baselines in both tasks. Beyond performance, we show that learned embeddings and codebook structure capture interpretable acoustic units aligned with dolphin whistle categories and possibly sub-whistle structure, enabling fine-grained analysis of communication patterns. Our findings demonstrate how SSL can serve as both a model and a scientific tool to explore hypotheses in animal communication research.

自监督学习动物通信语音分析海豚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。