MIDI-GPT可精准控制多轨音乐生成,支持风格、乐器等条件约束。
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
- 采用轨道分列序列表示法,实现多轨音乐的可控生成
- 生成音乐与训练数据无重复,且风格匹配度高
- 适合音乐创作、游戏配乐及工业级音乐工具开发
我们提出并发布MIDI-GPT,一个基于Transformer架构的生成系统,专为计算机辅助多轨音乐创作设计。MIDI-GPT支持在轨和小节级别进行音乐内容填充,并能根据乐器类型、音乐风格、音符密度、和声复杂度和音符时长等属性进行条件生成。为整合这些特征,我们采用一种替代性表示:为每条轨道创建时间有序的音乐事件序列,并将多个轨道拼接成单一序列,而非将不同轨道的事件交错排列。我们还提出一种增强表现力的变体表示方法。实验结果表明,MIDI-GPT能持续避免复制训练数据中的音乐内容,生成风格与训练集一致的音乐,且属性控制可有效施加各类约束。我们还展示了MIDI-GPT在实际应用中的潜力,包括与产业伙伴合作探索其在商业产品中的集成与评估,以及使用该模型创作的多件艺术作品。
原文摘要 · Abstract (English)
We present and release MIDI-GPT, a generative system based on the Transformer architecture that is designed for computer-assisted music composition workflows. MIDI-GPT supports the infilling of musical material at the track and bar level, and can condition generation on attributes including: instrument type, musical style, note density, polyphony level, and note duration. In order to integrate these features, we employ an alternative representation for musical material, creating a time-ordered sequence of musical events for each track and concatenating several tracks into a single sequence, rather than using a single time-ordered sequence where the musical events corresponding to different tracks are interleaved. We also propose a variation of our representation allowing for expressiveness. We present experimental results that demonstrate that MIDI-GPT is able to consistently avoid duplicating the musical material it was trained on, generate music that is stylistically similar to the training dataset, and that attribute controls allow enforcing various constraints on the generated material. We also outline several real-world applications of MIDI-GPT, including collaborations with industry partners that explore the integration and evaluation of MIDI-GPT into commercial products, as well as several artistic works produced using it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。