arXiv:2411.07186cs.SDcs.AI2024-11被引 59

首个专为生物声学设计的音视频基础模型,能零样本识别未知物种叫声。

NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics

  • 构建跨领域文本-音频对数据集,融合生物声学、语音与音乐数据训练
  • 在新基准BEANS-Zero上实现零样本物种分类新纪录,泛化能力突出
  • 开源模型、数据与代码,助力生态保护与动物行为研究

大型语言模型通过文本与音频输入,在语音、音乐和通用音频任务中已达到顶尖性能,并展现出对未见任务的涌现能力。然而其在生物声学领域的潜力尚未充分释放,如在大规模录音中检测动物叫声、分类稀有濒危物种、标注情境与行为等关键任务。本文提出NatureLM-audio,首个专为生物声学设计的音视频基础模型。训练数据包含精心筛选的跨领域文本-音频对,涵盖生物声学、语音与音乐,以缓解标注数据稀缺问题。实验表明,模型可成功将音乐与语音中学到的表征迁移至生物声学任务,展现出对未见类群和任务的良好泛化能力。在新提出的BEANS-Zero基准上,模型在多个生物声学任务中达到新最优表现,包括零样本物种分类。为推动该领域发展,我们公开模型权重、基准数据及训练与数据生成的开源代码。

原文摘要 · Abstract (English)

Large language models (LLMs) prompted with text and audio have achieved state-of-the-art performance across various auditory tasks, including speech, music, and general audio, showing emergent abilities on unseen tasks. However, their potential has yet to be fully demonstrated in bioacoustics tasks, such as detecting animal vocalizations in large recordings, classifying rare and endangered species, and labeling context and behavior -- tasks that are crucial for conservation, biodiversity monitoring, and animal behavior studies. In this work, we present NatureLM-audio, the first audio-language foundation model specifically designed for bioacoustics. Our training dataset consists of carefully curated text-audio pairs spanning bioacoustics, speech, and music, designed to address the field's limited availability of annotated data. We demonstrate successful transfer of learned representations from music and speech to bioacoustics, and our model shows promising generalization to unseen taxa and tasks. We evaluate NatureLM-audio on a novel benchmark (BEANS-Zero) and it sets a new state of the art on several bioacoustics tasks, including zero-shot classification of unseen species. To advance bioacoustics research, we release our model weights, benchmark data, and open-source the code for training and benchmark data generation and model training.

生物声学音视频模型零样本学习生态保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。