用生物分类结构提升动物声音识别与特征推断能力
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference
- 基于物种分类层级构建音视频-文本对齐模型
- 在6823种动物声音上实现未见物种的准确识别
- 可直接从叫声推断生态特征,适合生态监测研究
动物鸣叫为野生动物评估提供了关键信息,尤其在森林等复杂环境中有助于物种识别与生态监测。深度学习虽已实现基于叫声的自动物种分类,但对训练中未见过的物种仍具挑战。为此,我们提出AnimalCLAP,一种融合生物分类信息的语言-音频预训练框架,包含新数据集与模型。该数据集收录4,225小时录音,覆盖6,823个物种,并标注22项生态特征。AnimalCLAP模型通过引入分类层级对齐音频与文本表征,显著提升未见物种的识别能力。实验表明,该模型能有效从叫声中推断物种的生态与生物学特征,性能优于CLAP。数据集、代码与模型将公开发布于https://dahlian00.github.io/AnimalCLAP_Page/。
原文摘要 · Abstract (English)
Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have enabled automatic species classification from their vocalizations. However, classifying species unseen during training remains challenging. To address this limitation, we introduce AnimalCLAP, a taxonomy-aware language-audio framework comprising a new dataset and model that incorporate hierarchical biological information. Specifically, our vocalization dataset consists of 4,225 hours of recordings covering 6,823 species, annotated with 22 ecological traits. The AnimalCLAP model is trained on this dataset to align audio and textual representations using taxonomic structures, improving the recognition of unseen species. We demonstrate that our proposed model effectively infers ecological and biological attributes of species directly from their vocalizations, achieving superior performance compared to CLAP. Our dataset, code, and models will be publicly available at https://dahlian00.github.io/AnimalCLAP_Page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。