arXiv:2410.15573cs.SDcs.AI2024-10被引 17

开源音乐理解基准与模型,解决多模态训练数据不足问题

OpenMU: Your Swiss Army Knife for Music Understanding

  • 构建大规模音乐理解基准OpenMU-Bench,整合现有数据并新增标注
  • 训练的OpenMU模型在歌词理解与音乐工具使用任务上优于基线
  • 适合音乐科技研究者与创意音乐生产者使用

我们提出OpenMU-Bench,一个大规模基准套件,旨在缓解多模态语言模型训练中音乐数据稀缺的问题。通过利用现有数据集并自举新增标注,OpenMU-Bench扩展了音乐理解范畴,涵盖歌词理解与音乐工具使用。基于该基准,我们训练了音乐理解模型OpenMU,通过大量消融实验验证其性能,结果表明OpenMU优于如MU-Llama等基线模型。OpenMU与OpenMU-Bench均已开源,以促进音乐理解研究,并提升创意音乐生产效率。

原文摘要 · Abstract (English)

We present OpenMU-Bench, a large-scale benchmark suite for addressing the data scarcity issue in training multimodal language models to understand music. To construct OpenMU-Bench, we leveraged existing datasets and bootstrapped new annotations. OpenMU-Bench also broadens the scope of music understanding by including lyrics understanding and music tool usage. Using OpenMU-Bench, we trained our music understanding model, OpenMU, with extensive ablations, demonstrating that OpenMU outperforms baseline models such as MU-Llama. Both OpenMU and OpenMU-Bench are open-sourced to facilitate future research in music understanding and to enhance creative music production efficiency.

音乐理解多模态开源数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。