用多模态数据训练音乐理解生成模型,效果优于现有方法。
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
- 基于167.69小时多模态数据,融合音视频图文预训练编码器。
- 在4项任务中超越当前最优模型,尤其在文本生成音乐上表现突出。
- 适合研究音乐生成、跨模态理解的学者与开发者参考。
大型语言模型在文本、语音、图像和视频领域已取得显著进展,但多模态音乐理解与生成仍因缺乏高质量标注数据而研究不足。为此,我们构建了一个包含167.69小时多模态数据的数据集,涵盖文本、图像、视频及音乐标注。基于该数据集,我们提出MuMu-LLaMA模型,利用预训练的音乐、图像和视频编码器。音乐生成方面,整合AudioLDM 2与MusicGen。在音乐理解、文本到音乐生成、基于提示的音乐编辑及多模态音乐生成四项任务上的评估表明,MuMu-LLaMA优于当前最优模型,展现出在多模态音乐应用中的巨大潜力。
原文摘要 · Abstract (English)
Research on large language models has advanced significantly across text, speech, images, and videos. However, multi-modal music understanding and generation remain underexplored due to the lack of well-annotated datasets. To address this, we introduce a dataset with 167.69 hours of multi-modal data, including text, images, videos, and music annotations. Based on this dataset, we propose MuMu-LLaMA, a model that leverages pre-trained encoders for music, images, and videos. For music generation, we integrate AudioLDM 2 and MusicGen. Our evaluation across four tasks--music understanding, text-to-music generation, prompt-based music editing, and multi-modal music generation--demonstrates that MuMu-LLaMA outperforms state-of-the-art models, showing its potential for multi-modal music applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。