arXiv:2508.01178cs.SDcs.AI2025-08

打造统一音乐理解模型,兼顾旋律与歌词的综合分析。

Advancing the Foundation Model for Music Understanding

  • 设计新架构联合处理音乐和歌词信息
  • 在多任务数据集上训练,表现超越现有模型
  • 适合需要跨模态音乐分析的研究者

音乐信息检索(MIR)领域存在模型碎片化问题,专用模型仅擅长单一任务。本文提出统一的基础模型MuFun,实现对音乐的全面理解。该模型采用新型架构,联合处理乐器与歌词内容,并在覆盖多种任务的大规模数据集上进行训练,包括流派分类、音乐标签识别和问答等。为支持稳健评估,我们还提出了新的多维度音乐理解基准MuCUE(Music Comprehensive Understanding Evaluation)。实验表明,该模型在MuCUE各项任务中显著优于现有音频大语言模型,展现出顶尖的性能与泛化能力。

原文摘要 · Abstract (English)

The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified foundation model named MuFun for holistic music understanding. Our model features a novel architecture that jointly processes instrumental and lyrical content, and is trained on a large-scale dataset covering diverse tasks such as genre classification, music tagging, and question answering. To facilitate robust evaluation, we also propose a new benchmark for multi-faceted music understanding called MuCUE (Music Comprehensive Understanding Evaluation). Experiments show our model significantly outperforms existing audio large language models across the MuCUE tasks, demonstrating its state-of-the-art effectiveness and generalization ability.

音乐理解基础模型多模态音频LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。