开源音乐理解基准与模型,解决多模态训练数据不足问题
OpenMU: Your Swiss Army Knife for Music Understanding
- 构建大规模音乐理解基准OpenMU-Bench,整合现有数据并新增标注
- 训练的OpenMU模型在歌词理解与音乐工具使用任务上优于基线
- 适合音乐科技研究者与创意音乐生产者使用
我们提出OpenMU-Bench,一个大规模基准套件,旨在缓解多模态语言模型训练中音乐数据稀缺的问题。通过利用现有数据集并自举新增标注,OpenMU-Bench扩展了音乐理解范畴,涵盖歌词理解与音乐工具使用。基于该基准,我们训练了音乐理解模型OpenMU,通过大量消融实验验证其性能,结果表明OpenMU优于如MU-Llama等基线模型。OpenMU与OpenMU-Bench均已开源,以促进音乐理解研究,并提升创意音乐生产效率。
原文摘要 · Abstract (English)
We present OpenMU-Bench, a large-scale benchmark suite for addressing the data scarcity issue in training multimodal language models to understand music. To construct OpenMU-Bench, we leveraged existing datasets and bootstrapped new annotations. OpenMU-Bench also broadens the scope of music understanding by including lyrics understanding and music tool usage. Using OpenMU-Bench, we trained our music understanding model, OpenMU, with extensive ablations, demonstrating that OpenMU outperforms baseline models such as MU-Llama. Both OpenMU and OpenMU-Bench are open-sourced to facilitate future research in music understanding and to enhance creative music production efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。