首个可理解乐谱的指令模型,让AI读懂五线谱
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
- 用双阶段流程对齐乐谱编码器与语言模型
- 在巨量乐谱数据上训练,生成效果优于传统转录方法
- 适合音乐智能、作曲辅助等需要精准乐理理解的场景
近期多模态大语言模型在音频音乐理解方面取得显著进展,但符号化音乐——这一音乐结构的基础表达形式——仍未被充分探索。本文提出 MIDI-LLaMA,首个面向符号化音乐理解的指令跟随多模态大模型。通过两阶段流程(特征对齐与指令微调),将 MIDI 编码器 MusicBERT 与 Llama-3-8B 对齐。为支持训练,设计可扩展的标注流程,对 GiantMIDI-Piano 数据集进行细粒度元数据标注,构建了一个 MIDI-文本数据集。相较于在相同指令微调流程下使用 ABC 符号转录作为输入的基线模型,MIDI-LLaMA 在音乐描述生成和问答语义对齐任务中表现更优。人工评估进一步验证其在音乐理解、情感识别、创意生成及整体偏好上的优势。结果表明,将符号化音乐引入大语言模型能显著提升其音乐理解能力。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this work, we introduce MIDI-LLaMA, the first instruction-following MLLM for symbolic music understanding. Our approach aligns the MIDI encoder MusicBERT and Llama-3-8B via a two-stage pipeline comprising feature alignment and instruction tuning. To support training, we design a scalable annotation pipeline that annotates GiantMIDI-Piano with fine-grained metadata, resulting in a MIDI-text dataset. Compared with the baseline trained on converting MIDI into ABC notation under the same instruction-tuning procedure, MIDI-LLaMA substantially outperforms in captioning and semantic alignment in question answering. Human evaluation further confirms the advantages of MIDI-LLaMA in music understanding, emotion recognition, creativity, and overall preference. These findings demonstrate that incorporating symbolic music into large language models enhances their capacity for musical understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。