用多模态数据训练音乐理解模型,提升识别与分析能力。
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
- 通过图文音视频四模态指令微调,融合多种音乐信息。
- 在六个任务中达当前最优,显著优于单模态方法。
- 适合研究音乐生成、跨模态理解的开发者使用。
近期音乐大语言模型在音乐理解任务上取得进展,主要依赖音乐与文本输入的结合。然而,图像、视频及文本音乐特征等额外模态的潜力尚未被挖掘。为此,我们提出 DeepResonance,一个基于四模态指令微调的多模态音乐理解大模型,利用音乐、文本、图像和视频对齐数据进行训练。构建了 Music4way-MI2T、Music4way-MV2T 与 Music4way-Any2T 三个四模态训练与评估数据集,支持视觉与文本音乐特征融合。引入多采样 ImageBind 嵌入与预融合 Transformer 模块,增强模态间融合效果。模型在六项音乐理解任务中均达到领先性能,验证了辅助模态与结构设计的优势。代码、模型与数据已开源:github.com/sony/DeepResonance。
原文摘要 · Abstract (English)
Recent advancements in music large language models (LLMs) have significantly improved music understanding tasks, which involve the model's ability to analyze and interpret various musical elements. These improvements primarily focused on integrating both music and text inputs. However, the potential of incorporating additional modalities such as images, videos and textual music features to enhance music understanding remains unexplored. To bridge this gap, we propose DeepResonance, a multimodal music understanding LLM fine-tuned via multi-way instruction tuning with multi-way aligned music, text, image, and video data. To this end, we construct Music4way-MI2T, Music4way-MV2T, and Music4way-Any2T, three 4-way training and evaluation datasets designed to enable DeepResonance to integrate both visual and textual music feature content. We also introduce multi-sampled ImageBind embeddings and a pre-LLM fusion Transformer to enhance modality fusion prior to input into text LLMs, tailoring for multi-way instruction tuning. Our model achieves state-of-the-art performances across six music understanding tasks, highlighting the benefits of the auxiliary modalities and the structural superiority of DeepResonance. We open-source the codes, models and datasets we constructed: github.com/sony/DeepResonance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。