arXiv:2412.20146eess.AScs.SD2024-12被引 2

用自监督解耦学习提取整首鸟鸣嵌入,提升生物声学分析效率。

Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning

  • 设计双编码器模型,分离鸟鸣的通用与判别特征
  • 在大山雀数据集上聚类效果优于预训练模型和普通VAE
  • 可压缩嵌入维度并解释特征解耦机制,适合生态监测应用

本文提出一种基于解耦表征学习(DRL)的方法,从整首鸟鸣中提取语音嵌入。鸟鸣嵌入对大规模生物声学任务至关重要,现有自监督方法如变分自编码器(VAE)已在音符或音节级别取得成效。为将处理层级扩展至整首歌曲,本文将每段鸣唱视为泛化与判别性部分,采用双编码器分别学习这两部分。该方法在大山雀(Great Tits)数据集上通过聚类性能评估,结果优于对比的预训练模型和原始VAE。最后,论文分析了嵌入中的信息成分,进一步压缩其维度,并解释了鸟鸣表征的解耦特性。

原文摘要 · Abstract (English)

This paper addresses the extraction of the bird vocalization embedding from the whole song level using disentangled representation learning (DRL). Bird vocalization embeddings are necessary for large-scale bioacoustic tasks, and self-supervised methods such as Variational Autoencoder (VAE) have shown their performance in extracting such low-dimensional embeddings from vocalization segments on the note or syllable level. To extend the processing level to the entire song instead of cutting into segments, this paper regards each vocalization as the generalized and discriminative part and uses two encoders to learn these two parts. The proposed method is evaluated on the Great Tits dataset according to the clustering performance, and the results outperform the compared pre-trained models and vanilla VAE. Finally, this paper analyzes the informative part of the embedding, further compresses its dimension, and explains the disentangled performance of bird vocalizations.

鸟鸣识别自监督学习嵌入表示解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。