开源音频大模型突破多模态理解极限,支持10分钟长音频推理
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

- 统一编码器融合语音、声音、音乐三模态表征学习
- 支持10分钟长音频理解与多轮对话,性能超越闭源模型
- 适合音频理解、智能助手、内容分析等场景研究者使用
我们提出Audio Flamingo 3(AF3),一个完全开源的前沿大型音频语言模型,显著提升对语音、声音和音乐的推理与理解能力。AF3引入:(i) AF-Whisper,一种通过新策略联合学习三类模态(语音、声音、音乐)表征的统一音频编码器;(ii) 灵活的按需思考机制,支持链式思维推理;(iii) 多轮多音频对话;(iv) 最长10分钟的音频理解与推理(含语音);(v) 语音到语音交互。为实现这些能力,我们构建了多个大规模训练数据集,包括AudioSkills-XL、LongAudio-XL、AF-Think和AF-Chat,并采用五阶段课程式训练策略。仅使用开源音频数据训练,AF3在20+个(长音频)理解与推理基准上达到新SOTA,超越大量使用更大规模数据训练的开源及闭源模型。
原文摘要 · Abstract (English)
We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。