arXiv:2410.00344cs.SDcs.LG2024-10

用大语言模型让文本生成音乐更长更有序。

Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces

  • 把文本到音乐模型和大语言模型结合,生成有结构的长音乐
  • 可生成2.5分钟长、高度连贯且组织性强的音乐
  • 适合需要长篇结构性音乐创作的场景

基于Transformer的近期音乐生成方法上下文窗口最多可达一分钟,超出该范围后生成的音乐大多无结构。在更长上下文窗口下,从音乐数据中学习长尺度结构是一个极具挑战性的问题。本文提出将文本到音乐模型与大语言模型集成,以生成具有音乐形式的作品。论文讨论了实现这一集成所面临的关键挑战及解决方案。实验结果表明,该方法可生成长达2.5分钟、高度结构化、强组织性且连贯的音乐。

原文摘要 · Abstract (English)

Recent music generation methods based on transformers have a context window of up to a minute. The music generated by these methods is largely unstructured beyond the context window. With a longer context window, learning long-scale structures from musical data is a prohibitively challenging problem. This paper proposes integrating a text-to-music model with a large language model to generate music with form. The papers discusses the solutions to the challenges of such integration. The experimental results show that the proposed method can generate 2.5-minute-long music that is highly structured, strongly organized, and cohesive.

音乐生成长序列语言模型结构化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。