arXiv:2604.10905cs.SDcs.AI2026-04被引 14

Audio Flamingo Next 能理解30分钟长音频,支持语音、环境音和音乐的复杂推理。

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

论文配图:Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
图 1 · 摘自论文原文
  • 采用时序思维链,将推理步骤与音频时间戳对齐
  • 训练数据超100万小时,覆盖长音频与复杂任务
  • 开源三版本模型,适用于指令、推理与字幕生成

我们提出 Audio Flamingo Next(AF-Next),Audio Flamingo 系列中下一代最强的大型音频语言模型,旨在提升对语音、环境声音和音乐的理解与推理能力。相比 Audio Flamingo 3,AF-Next 引入:(i) 更强的基础音频语言模型,在多样化音频理解任务中显著提升准确率;(ii) 可扩展的数据构建策略,构建大规模音频理解与推理数据集,超越现有学术基准;(iii) 支持长达30分钟的复杂音频输入;(iv) 时序音频思维链(Temporal Audio Chain-of-Thought),通过显式关联中间推理步骤与音频时间戳,实现细粒度时空对齐与可解释性。为实现这些能力,我们首先系统分析 Audio Flamingo 3 的关键短板,随后构建并扩展超过100万小时的新数据集,涵盖 AudioSkills-XL、LongAudio-XL、AF-Think 和 AF-Chat。AF-Next 采用分阶段课程学习策略进行训练,覆盖预训练、中段训练与后训练。在20个音频理解与推理基准上,包括挑战性的长音频任务,结果表明其性能远超同等规模的开放模型,且常优于甚至超越更大规模的开放与闭源模型。此外,其在真实场景中表现出强实用性,并能良好泛化至未见任务,体现优异鲁棒性与泛化能力。我们开源了三个版本的 AF-Next 模型,包括 AF-Next-Instruct、AF-Next-Think 与 AF-Next-Captioner。

原文摘要 · Abstract (English)

We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.

音频理解长音频多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。