arXiv:2409.08673cs.SDcs.LG2024-09被引 2

用分层对比学习提升动物声音辨识精度

Acoustic identification of individual animals with hierarchical contrastive learning

  • 将个体动物声纹识别建模为分层多标签分类任务
  • 分层嵌入使个体与物种层级识别准确率均提升
  • 适合需要区分同种动物个体的研究场景

动物声音个体识别(AIID)与基于音频的物种分类密切相关,但需在同种内区分个体,要求更细粒度。本文将AIID建模为分层多标签分类任务,提出分层感知损失函数,以学习保持物种与分类阶元间层级关系的鲁棒嵌入表示。实验表明,分层嵌入不仅提升了个体层级的识别准确率,也提高了更高分类层级的表现,有效保留了学习表示中的层级结构。与非分层模型对比,验证了在嵌入空间中强制层级结构的优势。此外,方法在新个体类别分类(开集场景)上表现良好,展现出在开放环境下的应用潜力。

原文摘要 · Abstract (English)

Acoustic identification of individual animals (AIID) is closely related to audio-based species classification but requires a finer level of detail to distinguish between individual animals within the same species. In this work, we frame AIID as a hierarchical multi-label classification task and propose the use of hierarchy-aware loss functions to learn robust representations of individual identities that maintain the hierarchical relationships among species and taxa. Our results demonstrate that hierarchical embeddings not only enhance identification accuracy at the individual level but also at higher taxonomic levels, effectively preserving the hierarchical structure in the learned representations. By comparing our approach with non-hierarchical models, we highlight the advantage of enforcing this structure in the embedding space. Additionally, we extend the evaluation to the classification of novel individual classes, demonstrating the potential of our method in open-set classification scenarios.

声纹识别分层学习动物识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。