arXiv:2511.10289eess.AScs.CL2025-11被引 29

构建首个能深度理解音乐结构与文化的通用音频语言模型

Music Flamingo: Scaling Music Understanding in Audio Language Models

  • 基于多阶段标注构建大规模音乐理解数据集MF-Skills
  • 在10+评测中达顶尖性能,实现从表面识别到深层推理跃迁
  • 适合音乐智能、跨文化理解等研究者参考

我们提出Music Flamingo,一种新型大型音频-语言模型,旨在提升基础音频模型对音乐(包括歌曲)的理解能力。尽管音频-语言研究进展迅速,但音乐因其动态性、层次性和信息密集性仍具挑战。现有模型受限于高质量音乐数据和标注的稀缺,仅能生成短句级描述,回答表层问题,且跨文化泛化能力有限。为此,我们构建了MF-Skills——一个通过多阶段流水线标注的大规模数据集,包含涵盖和声、结构、音色、歌词及文化背景的丰富描述与问答对。我们在增强版Audio Flamingo 3基础上微调,并强化多项音乐理解技能。为提升推理能力,引入后训练策略:先以基于音乐理论的MF-Think链式思维数据集冷启动,再通过基于GRPO的强化学习与定制奖励优化。Music Flamingo在10多个音乐理解与推理基准上达到领先水平,展现出通用性与音乐智能。该工作不仅提供实证突破,更确立了向人类级音乐感知演进的新标准。

原文摘要 · Abstract (English)

We introduce Music Flamingo, a novel large audio-language model designed to advance music (including song) understanding in foundational audio models. While audio-language research has progressed rapidly, music remains challenging due to its dynamic, layered, and information-dense nature. Progress has been further limited by the difficulty of scaling open audio understanding models, primarily because of the scarcity of high-quality music data and annotations. As a result, prior models are restricted to producing short, high-level captions, answering only surface-level questions, and showing limited generalization across diverse musical cultures. To address these challenges, we curate MF-Skills, a large-scale dataset labeled through a multi-stage pipeline that yields rich captions and question-answer pairs covering harmony, structure, timbre, lyrics, and cultural context. We fine-tune an enhanced Audio Flamingo 3 backbone on MF-Skills and further strengthen multiple skills relevant to music understanding. To improve the model's reasoning abilities, we introduce a post-training recipe: we first cold-start with MF-Think, a novel chain-of-thought dataset grounded in music theory, followed by GRPO-based reinforcement learning with custom rewards. Music Flamingo achieves state-of-the-art results across 10+ benchmarks for music understanding and reasoning, establishing itself as a generalist and musically intelligent audio-language model. Beyond strong empirical results, Music Flamingo sets a new standard for advanced music understanding by demonstrating how models can move from surface-level recognition toward layered, human-like perception of songs. We believe this work provides both a benchmark and a foundation for the community to build the next generation of models that engage with music as meaningfully as humans do.

音乐理解音频语言模型多模态链式思维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。