arXiv:2410.02084cs.SDeess.AS2024-10中稿 · ISMIR 2025被引 13

用大模型生成自然语言描述,让音乐自动生成更自由可控。

Generating Symbolic Music from Natural Language Prompts using an LLM-Enhanced Dataset

  • 用大模型将音乐元数据转为自然语言描述,构建增强数据集
  • 基于伪文本提示的生成模型在听觉测试中优于基线
  • 支持自由描述词控制,适合想用说话方式创作音乐的人

近年来,音频域的文本到音乐生成模型依赖大量文本-音频对进行训练。但符号音乐领域的可控生成进展较慢,主要因缺乏大规模带丰富元数据和描述的符号音乐数据集。本文提出MetaScore,一个包含96.3万首乐谱与丰富元数据(包括用户自由标注标签)的数据集,数据源自在线音乐论坛。我们利用预训练大语言模型(LLM)从乐谱元数据标签生成伪自然语言描述。基于该增强数据集,训练了一个文本条件的符号音乐生成模型,可依据伪描述生成音乐,实现对乐器、风格、作曲家、复杂度等自由描述词的控制。此外,还训练了标签条件系统,支持MetaScore中预定义标签。实验表明,两种模型在听觉测试中均优于基线模型。尽管同期工作Text2MIDI也支持自由文本输入,我们的模型性能相当。且文本生成模型提供更自然的交互界面,允许用户输入自由语言提示。

原文摘要 · Abstract (English)

Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training. However, symbolic-domain controllable music generation has lagged behind partly due to the lack of a large-scale symbolic music dataset with extensive metadata and captions. In this work, we present MetaScore, a new dataset consisting of 963K musical scores paired with rich metadata, including free-form user-annotated tags, collected from an online music forum. To approach text-to-music generation, We employ a pretrained large language model (LLM) to generate pseudo-natural language captions for music from its metadata tags. With the LLM-enhanced MetaScore, we train a text-conditioned music generation model that learns to generate symbolic music from the pseudo captions, allowing control of instruments, genre, composer, complexity and other free-form music descriptors. In addition, we train a tag-conditioned system that supports a predefined set of tags available in MetaScore. Our experimental results show that both the proposed text-to-music and tags-to-music models outperform a baseline text-to-music model in a listening test. While a concurrent work Text2MIDI also supports free-form text input, our models achieve comparable performance. Moreover, the text-to-music system offers a more natural interface than the tags-to-music model, as it allows users to provide free-form natural language prompts.

文本生成音乐符号音乐大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。