大模型能从文字生成音乐,揭示其隐含的音乐结构理解能力
Large Language Models' Internal Perception of Symbolic Music
- 用文本描述生成无训练数据的MIDI音乐文件
- 基于生成数据训练的模型达到可比分类与补全性能
- 适合研究大模型对符号化音乐的隐式建模能力
大型语言模型(LLMs)在自然语言字符串关系建模上表现优异,并展现出向代码、数学等符号领域扩展的潜力。然而,其对符号音乐的隐式建模程度仍不明确。本文通过文本提示生成包含风格与流派组合的符号音乐数据,评估其在识别与生成任务中的效用。我们构建了一个不依赖显式音乐训练的LLM生成MIDI数据集,并在此基础上训练神经网络,完成流派与风格分类及旋律补全任务,与现有模型对比性能。结果表明,LLMs能从文本中推断出基本的音乐结构与时间关系,既展示了其隐式编码音乐模式的潜力,也暴露了因缺乏显式音乐背景导致的局限性,为理解其符号音乐生成能力提供了洞见。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at modeling relationships between strings in natural language and have shown promise in extending to other symbolic domains like coding or mathematics. However, the extent to which they implicitly model symbolic music remains underexplored. This paper investigates how LLMs represent musical concepts by generating symbolic music data from textual prompts describing combinations of genres and styles, and evaluating their utility through recognition and generation tasks. We produce a dataset of LLM-generated MIDI files without relying on explicit musical training. We then train neural networks entirely on this LLM-generated MIDI dataset and perform genre and style classification as well as melody completion, benchmarking their performance against established models. Our results demonstrate that LLMs can infer rudimentary musical structures and temporal relationships from text, highlighting both their potential to implicitly encode musical patterns and their limitations due to a lack of explicit musical context, shedding light on their generative capabilities for symbolic music.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。