用元数据控制生成4小节多轨音乐,灵活又保质。
Flexible Control in Symbolic Music Generation via Musical Metadata
- 用自回归模型,输入音乐元数据生成4小节多轨MIDI。
- 随机丢弃元数据训练,提升生成控制灵活性。
- 适合需要精准控制旋律主题的作曲者使用。
本文提出一种符号化音乐生成方法,聚焦于生成以短音乐动机为核心的叙事性主题。采用自回归模型,以音乐元数据为输入,生成4小节多轨MIDI序列。训练时随机丢弃元数据中的部分标记,以确保生成过程具备灵活的控制能力。用户可自由选择输入类型,同时保持生成质量,显著提升作曲灵活性。通过模型容量、音乐保真度、多样性及可控性等实验验证策略有效性。进一步扩大模型规模,并在主观测试中与其它音乐生成模型对比,结果表明该方法在控制能力和音乐质量上均具优势。演示视频见:https://www.youtube.com/watch?v=-0drPrFJdMQ。
原文摘要 · Abstract (English)
In this work, we introduce the demonstration of symbolic music generation, focusing on providing short musical motifs that serve as the central theme of the narrative. For the generation, we adopt an autoregressive model which takes musical metadata as inputs and generates 4 bars of multitrack MIDI sequences. During training, we randomly drop tokens from the musical metadata to guarantee flexible control. It provides users with the freedom to select input types while maintaining generative performance, enabling greater flexibility in music composition. We validate the effectiveness of the strategy through experiments in terms of model capacity, musical fidelity, diversity, and controllability. Additionally, we scale up the model and compare it with other music generation model through a subjective test. Our results indicate its superiority in both control and music quality. We provide a URL link https://www.youtube.com/watch?v=-0drPrFJdMQ to our demonstration video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。