arXiv:2511.10190cs.LG2025-11中稿 · NeurIPS被引 2

用离散声学标记序列捕捉动物鸣叫的时序结构,提升分类性能。

Towards Leveraging Sequential Structure in Animal Vocalizations

  • 通过向量量化提取鸣叫的离散声学标记序列,保留时序信息。
  • 在4个数据集上,标记序列可有效区分叫声类型和发声个体。
  • 适合研究动物交流行为、语音建模或生物声学特征提取的学者。

动物鸣叫包含重要的时序结构信息,但现有计算生物声学研究通常对帧级特征沿时间轴求平均,忽略了子单元的顺序。本文探讨通过自监督语音模型表征的向量量化与Gumbel-Softmax向量量化生成的离散声学标记序列,能否有效捕捉并利用时序信息。基于HuBERT嵌入的标记序列对四组生物声学数据进行成对距离分析显示,其可区分叫声类型与发声个体。使用Levenshtein距离的k近邻分类实验表明,该方法在呼叫类型与发声者分类任务中表现良好,具有作为时序信息建模替代特征表示的潜力。

原文摘要 · Abstract (English)

Animal vocalizations contain sequential structures that carry important communicative information, yet most computational bioacoustics studies average the extracted frame-level features across the temporal axis, discarding the order of the sub-units within a vocalization. This paper investigates whether discrete acoustic token sequences, derived through vector quantization and gumbel-softmax vector quantization of extracted self-supervised speech model representations can effectively capture and leverage temporal information. To that end, pairwise distance analysis of token sequences generated from HuBERT embeddings shows that they can discriminate call-types and callers across four bioacoustics datasets. Sequence classification experiments using $k$-Nearest Neighbour with Levenshtein distance show that the vector-quantized token sequences yield reasonable call-type and caller classification performances, and hold promise as alternative feature representations towards leveraging sequential information in animal vocalizations.

声学建模时序结构动物交流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。