arXiv:2511.03942cs.SDcs.CL2025-11中稿 · International Soci…被引 4

用大模型提升文本生成多轨音乐质量,支持真实创作流程。

MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models

  • 将大模型文本能力扩展至音乐符号,分两阶段训练提升生成效果。
  • 在58名用户4002次生成中,零起点创作接受率超基线模型。
  • 适配实际作曲流程,支持旋律+和弦生成与补全任务。

我们提出MIDI-LLM,一种通过适配大语言模型(LLM)改进多轨文本到MIDI生成的方案。MIDI-LLM将LLM的文本词汇扩展为包含MIDI标记,并采用两阶段训练:(i) 在音乐相关文本和独立MIDI上进行单模态持续预训练;(ii) 在文本-MIDI对上进行多模态监督微调。基于Llama 3.2(1B)的实例在文本控制与音乐质量上均优于近期Text2midi模型,且可无缝集成vLLM等优化推理生态。为契合真实歌曲创作流程,我们在TheoryTab数据集上进一步微调以实现文本条件下的乐谱(旋律+和弦)生成与补全。全面消融实验验证了文本预训练、独立MIDI预训练与监督微调之间的协同效应。最后,在58名Hookpad Aria用户参与的真实创作场景下开展盲测,共生成4002条作品,结果表明我们的MIDI-LLM在无文本控制或无LLM预训练的基线模型中,零到一生成的接受率最高,证实其在人机协作创作中的有效性。

原文摘要 · Abstract (English)

We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an LLM's text vocabulary to include MIDI tokens and employs a two-stage training pipeline: (i) unimodal continued pretraining on music-adjacent text and standalone MIDIs, and (ii) multimodal supervised finetuning on text-MIDI pairs. Our instantiation of MIDI-LLM based on Llama 3.2 (1B) outperforms the recent Text2midi model in both text control and musical quality, and readily integrates with optimized inference ecosystems like vLLM. To align with real-world songwriting workflows, we further finetune our MIDI-LLM on the TheoryTab dataset for text-conditioned lead sheet (i.e., melody + chords) generation and infilling. A comprehensive ablation study validates the synergy between LLM text pretraining, standalone MIDI pretraining, and supervised text-to-MIDI finetuning. Finally, an in-the-wild blind user study conducted in a real-world creative workflow at scale with 58 Hookpad Aria users and 4,002 generated outputs demonstrates that our MIDI-LLM achieves the highest acceptance rate in zero-to-one lead sheet generation over baselines without text control or LLM pretraining, confirming its efficacy in human-AI music co-creation.

文本生成音乐生成大模型多轨音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。