arXiv:2503.03983cs.SDcs.CL2025-03ICML被引 179

小模型实现音频理解与推理新突破,支持长音频分析

Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities

  • 用定制CLAP模型+合成问答数据+分阶段训练提升音频推理能力
  • 30亿参数模型在20+基准上超越大模型表现
  • 首次构建长音频数据集,支持30秒至5分钟音频理解评测

理解和推理非语音声音与音乐对人类和人工智能体有效互动至关重要。本文提出Audio Flamingo 2(AF2),一种具备先进音频理解与推理能力的音频-语言模型。AF2采用(i)定制的CLAP模型,(ii)用于细粒度音频推理的合成音频问答数据,以及(iii)多阶段课程学习策略。仅使用30亿参数的小型语言模型,AF2在超过20个基准上达到领先性能,超越多个大型开源及专有模型。此外,首次将音频理解扩展至长音频段(30秒至5分钟),并提出LongAudio——一个大规模新型数据集,用于训练音频-语言模型进行长音频描述与问答任务。在LongAudio上微调AF2后,在自建的LongAudioBench(专家标注的长音频理解评测基准)上取得卓越表现。通过大量消融实验验证了方法的有效性。

原文摘要 · Abstract (English)

Understanding and reasoning over non-speech sounds and music are crucial for both humans and AI agents to interact effectively with their environments. In this paper, we introduce Audio Flamingo 2 (AF2), an Audio-Language Model (ALM) with advanced audio understanding and reasoning capabilities. AF2 leverages (i) a custom CLAP model, (ii) synthetic Audio QA data for fine-grained audio reasoning, and (iii) a multi-stage curriculum learning strategy. AF2 achieves state-of-the-art performance with only a 3B parameter small language model, surpassing large open-source and proprietary models across over 20 benchmarks. Next, for the first time, we extend audio understanding to long audio segments (30 secs to 5 mins) and propose LongAudio, a large and novel dataset for training ALMs on long audio captioning and question-answering tasks. Fine-tuning AF2 on LongAudio leads to exceptional performance on our proposed LongAudioBench, an expert annotated benchmark for evaluating ALMs on long audio understanding capabilities. We conduct extensive ablation studies to confirm the efficacy of our approach. Project Website: https://research.nvidia.com/labs/adlr/AF2/.

音频理解长音频多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。