用自动转录音频训练音乐生成模型,支持提示与约束控制。
SymPAC: Scalable Symbolic Music Generation With Prompts And Constraints
- 用预训练音乐信息检索模型从音频自动生成符号音乐序列
- 仅依赖自动转录数据即可训练出高质量符号音乐生成模型
- 支持提示词引导和有限状态机约束,适合创作场景
符号音乐生成进展相对滞后,部分原因在于符号训练数据稀缺。本文利用大规模音频音乐数据,通过预训练的音乐信息检索(MIR)模型(如转录、节拍检测、结构分析等)提取符号事件,并将其编码为标记序列。据我们所知,这是首个仅基于自动转录音频数据训练符号生成模型的成功范例。此外,为增强模型可控性,提出SymPAC(Symbolic Music Language Model with Prompting And Constrained Generation),其特点包括:(a) 编码时引入提示栏(prompt bars);(b) 推理时采用基于有限状态机(FSMs)的约束生成技术。实验表明该方法具有高度灵活性与可控性,对创作者和用户而言极具实用价值。
原文摘要 · Abstract (English)
Progress in the task of symbolic music generation may be lagging behind other tasks like audio and text generation, in part because of the scarcity of symbolic training data. In this paper, we leverage the greater scale of audio music data by applying pre-trained MIR models (for transcription, beat tracking, structure analysis, etc.) to extract symbolic events and encode them into token sequences. To the best of our knowledge, this work is the first to demonstrate the feasibility of training symbolic generation models solely from auto-transcribed audio data. Furthermore, to enhance the controllability of the trained model, we introduce SymPAC (Symbolic Music Language Model with Prompting And Constrained Generation), which is distinguished by using (a) prompt bars in encoding and (b) a technique called Constrained Generation via Finite State Machines (FSMs) during inference time. We show the flexibility and controllability of this approach, which may be critical in making music AI useful to creators and users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。